March 2021 Summaries
25 posts from Datadog
Filter
Month:
Year:
Post Summaries
Back to Blog
AWS users can now forward metrics from key AWS services to different endpoints, including Datadog, via Amazon Kinesis Data Firehose with low latency using CloudWatch Metric Streams. This feature allows for up to an 80 percent reduction in latency compared to using GetMetricData API calls. By integrating CloudWatch Metric Streams with Datadog, teams can obtain comprehensive visibility into their AWS services' health and performance through low-latency metrics and logs. The integration process involves setting up a Kinesis Data Firehose delivery stream and linking it to a CloudWatch Metric Stream for data ingestion from specified AWS services. This enables users to monitor key AWS services such as ELB, RDS, ElastiCache, and more with significantly lower latency, allowing them to identify issues quickly and assess the effectiveness of deployed fixes.
Mar 31, 2021
991 words in the original blog post.
Datadog Network Performance Monitoring (NPM) now supports Windows hosts, providing comprehensive visibility into hybrid and cloud environments with mixed operating systems. NPM enables users to visualize network traffic flow across endpoints and contextualize application and infrastructure issues in multi-OS environments. The Network Map displays the volume of bytes sent between services, while the Network Overview provides customizable views of network dependencies and their health status. By integrating NPM with logs, infrastructure metrics, and application traces, users can quickly pinpoint root causes of issues within Windows or multi-OS networks without switching contexts.
Mar 31, 2021
1,222 words in the original blog post.
Datadog is partnering with AWS to launch CloudWatch Metric Streams, a feature that allows users to forward metrics from key AWS services to different endpoints, including Datadog, via Amazon Data Firehose. This integration reduces latency by up to 80% compared to using GetMetricData API calls, providing low-latency metrics and logs for comprehensive monitoring of AWS services' health and performance. By setting up a CloudWatch Metric Stream with Datadog, users can gain faster visibility into key AWS services such as ELB, RDS, ElastiCache, and more, enabling them to quickly identify issues and respond to problems. This integration is particularly useful for monitoring high-traffic services like grocery delivery platforms during unprecedented demand. With the partnership, users can start sending metrics to Datadog for analysis and troubleshooting of their AWS infrastructure with reduced latency and improved visibility.
Mar 31, 2021
996 words in the original blog post.
Fastly is an edge cloud platform that offers various services such as content delivery network (CDN), image optimization, video streaming, cloud security, and load balancing. Datadog's Fastly integration allows users to visualize, analyze, and alert on Fastly metrics and logs in context with monitoring data from across their entire stack. This enables quick resolution of issues before they degrade the user experience. The integration provides comprehensive monitoring capabilities for key Fastly metrics like hit ratios, cache coverage, header size, error percentage, and HTTP client and server errors. It also allows users to pivot easily between Fastly metrics and related logs for better understanding and investigation of concerning activity. Additionally, Datadog's tag-based alerting system enables users to create sophisticated monitors for their Fastly services and collaborate efficiently during the troubleshooting process.
Mar 29, 2021
658 words in the original blog post.
Microsoft 365, previously known as Office 365, is a cloud-based suite of productivity tools used globally by over a million companies, necessitating effective monitoring to mitigate downtime and ensure optimal use. RapDev's integration with Datadog offers users a comprehensive way to track Microsoft 365 activities, usage, and licensing across its applications, such as Exchange and Teams, through robust dashboards and Synthetic tests. This integration allows users to visualize key performance metrics, conduct proactive Synthetic tests on user workflows, validate email functionality, and monitor license usage effectively, thereby enhancing operational efficiency and reducing unnecessary licensing costs. Available in the Datadog Marketplace, this integration provides organizations with the tools to quickly identify and resolve issues while offering insights into application-specific usage metrics. Users can experience the integration through a free two-week trial, facilitating a deeper understanding of its benefits for their business operations.
Mar 22, 2021
706 words in the original blog post.
SonarQube is a tool for static code analysis that integrates with CI pipelines to run quality checks on codebases as they change, helping ensure compliance, stability, and security. Datadog's SonarQube integration provides visibility into key metrics and logs, enabling real-time monitoring of code quality and server health. The integration allows users to visualize and analyze code metrics, collect and analyze SonarQube logs, alert on code-level security issues, and track the health of their SonarQube server. With Datadog's log processing pipeline, users can filter, sort, and search their logs to troubleshoot issues, and define alerts to notify on critical vulnerabilities, such as SQL injection. The integration also enables correlation with other CI pipeline metrics, providing a comprehensive view of code deployments and pipeline health.
Mar 18, 2021
781 words in the original blog post.
With Datadog, .NET developers can collect, visualize, and alert on key runtime metrics to troubleshoot bugs, detect resource inefficiencies, and optimize application performance. Monitoring first-chance exceptions helps identify unexpected errors that could degrade application performance, while detecting thread pool starvation optimizes multithreading tasks. Ensuring efficient garbage collection also contributes to optimal application performance by minimizing CPU usage and memory pressure. By visualizing .NET runtime metrics alongside other APM monitoring data, developers can gain a complete view of their applications in one place, enabling them to quickly identify issues and take corrective action.
Mar 17, 2021
670 words in the original blog post.
Datadog has introduced the ability to generate globally accurate aggregations of process metrics across any subset of applications and infrastructure. These metrics will be stored at full granularity for 15 months, providing rich historical context for exploring trends in data and troubleshooting performance degradations of infrastructure components. This feature allows users to create and manage process metrics, analyze historical trends in infrastructure load, leverage percentile aggregates to spot outlying processes, use processes alongside other telemetry data to identify the root cause of issues, and detect future issues more proactively with alerts and SLOs.
Mar 16, 2021
1,462 words in the original blog post.
VoltDB is an ACID-compliant in-memory relational database designed for real-time analytics and optimized for fast data processing. Datadog's new VoltDB integration offers full visibility into key metrics and logs, enabling users to monitor and alert on the health and performance of their databases. The dashboard provides insights into available memory, query latency, stored procedure performance, and table resource usage. It also includes a log stream for garbage collection, client authentication, and SQL execution events. Datadog's integration allows users to track table state, including row count and distribution, as well as monitor the performance of stored procedures. This helps ensure optimal database performance and provides valuable context for troubleshooting any issues that may arise.
Mar 11, 2021
550 words in the original blog post.
Datadog APM has introduced native tracing for AWS Lambda functions in Python and Node.js to provide deep visibility into serverless applications. The latest enhancements connect Lambda functions with AWS managed services all in one trace, allowing developers to effectively debug event-driven serverless applications by understanding where an issue occurred and how upstream and downstream services were involved. Datadog APM also tags function spans with additional information about incoming events for quick searching, filtering, and aggregating data when troubleshooting issues. This end-to-end visibility helps developers find and fix problems faster in their serverless applications.
Mar 09, 2021
847 words in the original blog post.
Datadog APM now connects Python and Node.js Lambda functions to AWS managed services in a single trace, allowing developers to understand where issues occurred and how upstream and downstream services were involved. This enables tracing of event-driven serverless applications across multiple components, including message queues, data streams, notification services, and more. The new feature provides additional tags for function spans, enabling quick search, filtering, and aggregation of data when troubleshooting issues. With this enhancement, Datadog brings distributed traces into the same view as infrastructure metrics and logs, providing detailed context around event-driven architectures. Developers can prioritize fixes more strategically by identifying the scope of an issue and its impact on end users. The new approach to distributed tracing embraces the complexities of modern serverless applications, giving end-to-end visibility into event-driven serverless applications to help find and fix issues faster.
Mar 09, 2021
861 words in the original blog post.
Microsoft Azure provides cloud computing services that enable organizations to deploy and manage web applications across various industries. As the usage of Azure-based applications expands, securing all cloud resources becomes increasingly complex. Azure platform logs record user activity within an Azure environment, including who performed an action, what was done, when it occurred, and where it took place. Monitoring these logs is crucial for maintaining the security of Azure assets and identifying potential malicious activities before they can spread throughout the system.
Azure uses Azure Active Directory (Azure AD) to manage identity and access management across all resources within an organization. The organizational hierarchy of Azure resource directory consists of four levels: management groups, subscriptions, resource groups, and resources. Each level acts as a hierarchy, with permissions configured for an entity at a higher level applying to all sub-resources within that entity.
Azure generates three categories of platform logs: Azure Active Directory reports, activity logs, and resource logs. Active Directory reports detail changes made in Azure AD and login activity, while activity logs record operations performed on an Azure resource, such as creating a virtual machine or editing the configuration of an Azure Pipeline. Resource logs capture operations within an existing Azure resource, like reads and writes to a vault in Azure Key Vault or to a database in Azure SQL Database.
To interpret Azure platform logs, it is essential to understand common information shared across all log types, such as the caller field (identity of the user or service that performed the logged action), category field (which helps determine the log type), and other fields like resource group, subscription ID, and operation name.
Monitoring key Azure platform logs can help detect potential vulnerabilities in an environment. Authentication logs provide a record of user activity, including login events, while resource-based logs focus on instances of resources with overly permissive access policies. By using a third-party log management solution like Datadog, organizations can gain a big-picture perspective of their Azure environment's activity and easily monitor these critical logs for potential threats.
To ship Azure platform logs to Datadog, it is recommended to use Event Hubs, which are distributed data streaming pipelines that handle the large volume of platform logs generated by an Azure environment. Once the logs are collected with Datadog, custom dashboards can be created to visualize log data for a comprehensive understanding of the activity in the Azure environment. Additionally, built-in Threat Detection Rules automatically watch the logs for potential malicious activities and notify users as soon as security and compliance issues occur.
Mar 05, 2021
3,171 words in the original blog post.
Microsoft Azure provides a suite of cloud computing services that allow organizations to deploy, manage, and monitor full-scale web applications. As the complexity of securing these applications increases, collecting and analyzing Azure platform logs becomes crucial for monitoring security and identifying potential threats. The logs are generated in three categories: Microsoft Entra ID reports detail changes made in Entra ID and login activity, Activity logs record operations performed on an Azure resource, such as creating a VM, Resource logs capture operations performed within an existing Azure resource. Each log type has unique fields that provide valuable information for tracking actions occurring in the environment. Understanding the organizational hierarchy of Azure resources is essential to properly interpreting and acting on those logs. Datadog can help organizations collect and monitor their Azure logs by providing automatic parsing and enrichment, cost-effective collection and archiving, built-in security and compliance analysis, and a user-friendly interface for visualizing log data and detecting security threats in real-time.
Mar 05, 2021
2,872 words in the original blog post.
The text discusses the implementation of wildcard-filtered metric queries in Datadog, a tool designed to enhance the efficiency of data monitoring in dynamic cloud infrastructures. By utilizing the wildcard asterisk (*) in queries, users can easily filter and track data across varying scopes without the need for repetitive and complex filtering processes. This method is particularly beneficial for managing environments with frequently changing assets, such as virtual machines and containers, as it allows for the automatic exclusion of irrelevant data like temporary file systems. Wildcard filtering can be combined with boolean syntax for more sophisticated queries, providing significant flexibility and adaptability in monitoring tasks. Additionally, this feature is available for immediate use by Datadog customers, along with a 14-day free trial for new users.
Mar 05, 2021
573 words in the original blog post.
NerdVision, a live debugging platform, has integrated with Datadog to provide users with enhanced visibility into their applications' performance. The integration allows developers to take snapshots of application states at runtime and visualize debugging trends without any changes to the source code. By sending all NerdVision data to Datadog, users can correlate code-level insights with monitoring data from their entire tech stack, reducing mean time to resolution (MTTR). The integration also provides an out-of-the-box dashboard for real-time monitoring of incoming tracepoints and logs, enabling developers to identify areas in need of improvement. Datadog Marketplace now offers the NerdVision integration, with plans to expand it further.
Mar 04, 2021
658 words in the original blog post.
NerdVision is a live debugging platform that enables users to take snapshots of their application's state at runtime, compatible with multiple programming languages and hosting environments. The new NerdVision integration with Datadog allows users to visualize debugging trends, reduce Mean Time To Recovery (MTTR), and correlate code-level insights with monitoring data from their entire tech stack. With the integration, users can take snapshots of their application's state at runtime using tracepoints and logs, which are then ingested into Datadog for visualization and analysis. The NerdVision toolbox provides flexible debugging tools that provide deep, code-level visibility into performance without any restarts or code changes, while also allowing users to monitor and analyze their data within the Datadog platform.
Mar 04, 2021
671 words in the original blog post.
Watchdog Insights is an AI-powered recommendation engine that augments incident investigations by highlighting parts of systems and applications with unusual warning signs. It helps users diagnose code-level issues more quickly, automatically analyzing logs to uncover noteworthy trends. The tool can help both small and large teams investigate complex issues involving logs from various sources such as applications, web servers, and infrastructure. By identifying log fields commonly associated with error messages or anomalies, Watchdog Insights provides the context needed for investigating infrastructure errors and finding suspect dependencies. It also helps determine which users are encountering problems in access logs, allowing teams to focus their investigations more efficiently.
Mar 02, 2021
974 words in the original blog post.
Juniper Networks offers various IT network and security devices such as routers, switches, access points, and firewalls. Datadog's Network Device Monitoring now supports Juniper devices, allowing users to collect metrics from their on-prem and virtual Juniper devices for optimal performance and health maintenance. The integration enables automatic recognition and collection of metrics from all Juniper devices alongside other vendor integrations like Cisco, Dell, and F5. Users can view these metrics in Datadog's out-of-the-box dashboard, which provides a centralized high-level view of the entire network's health and performance. The integration also supports viewing key Juniper metrics alongside other network infrastructure components. Alerts can be set to identify potential issues and initiate troubleshooting immediately. With this new integration and Network Device Monitoring, users can easily monitor and visualize metrics from Juniper devices within their on-prem or hybrid network alongside the rest of their IT infrastructure.
Mar 02, 2021
530 words in the original blog post.
Watchdog Insights is a recommendation engine that helps investigators identify parts of their systems and applications with an outsize proportion of warning signs, allowing them to diagnose issues more quickly. It surfaces error and latency outliers within request traces, analyzes logs to uncover noteworthy trends, and provides context around infrastructure errors. By automatically analyzing logs and identifying high-percentage fields in log searches, Watchdog Insights helps investigators focus on specific areas of their systems, reducing investigation time and improving efficiency. The tool can be used with existing workflows, making it suitable for small teams and big teams alike, and is part of a suite of features that provide clues to speed up investigations.
Mar 02, 2021
987 words in the original blog post.
Datadog has introduced a new integration with Juniper Networks to provide network device monitoring capabilities. This allows users to collect metrics from their on-prem and virtual Juniper devices, ensuring optimal performance and health of their IT infrastructure even at scale. The integration enables automatic recognition and collection of metrics, as well as tagging with device-level metadata for contextualization and filtering. Users can view Juniper metrics in the Datacenter Overview dashboard or Interface Performance dashboard, and set alerts to identify potential network issues before they spread. With this new integration, users can easily monitor and visualize their corporate IT infrastructure alongside other devices and systems.
Mar 02, 2021
538 words in the original blog post.
The guide provides a comprehensive walkthrough of using Datadog to monitor Amazon ECS and EKS applications running on AWS Fargate, detailing the integration of AWS environments with Datadog to collect and analyze metrics, logs, and traces. It outlines the process of enabling necessary integrations, deploying the Datadog Agent as a sidecar container, and configuring the system to visualize application performance through dashboards and alerts. The text emphasizes the importance of tagging for organizing data and understanding the health and performance of clusters, while also covering distributed tracing and log aggregation for troubleshooting. Additionally, it highlights the utility of Datadog's features, such as flame graphs and alerts, to identify and address performance issues in containerized infrastructures.
Mar 01, 2021
3,290 words in the original blog post.
In this article, we discuss how to monitor Amazon ECS and Amazon EKS clusters running on AWS Fargate using various tools such as Amazon CloudWatch, kubectl, Prometheus, and Fluent Bit. We cover collecting metrics from ECS on Fargate, collecting metrics from EKS on Fargate, and collecting logs from both ECS and EKS on Fargate. Additionally, we provide guidance on using CloudWatch Logs Insights to query and analyze collected logs. The article also highlights the use of Datadog for comprehensive monitoring data visualization and alerting in a single platform.
Mar 01, 2021
3,891 words in the original blog post.
AWS Fargate is a serverless container platform that allows users to run containers without managing the underlying infrastructure. It supports both Amazon Elastic Container Service (ECS) and Amazon Elastic Kubernetes Service (EKS). Fargate enables users to focus on delivering business value instead of capacity planning and operational issues. Monitoring Fargate clusters can help understand the performance of containerized applications and manage costs, as pricing is based on usage. Key metrics to monitor include memory utilization, CPU utilization, cluster state metrics, and other application-specific metrics.
Mar 01, 2021
3,477 words in the original blog post.
AWS Fargate provides a serverless container platform that allows developers to deploy and manage containerized workloads without provisioning or maintaining the underlying infrastructure. It integrates with Amazon Elastic Container Service (ECS) and Amazon Elastic Kubernetes Service (EKS), enabling users to focus on delivering business value instead of managing capacity planning and operational issues. Fargate pricing is based on usage, making it easier for developers to manage costs. To monitor Fargate clusters, users can track key metrics such as memory utilization, CPU utilization, and cluster state metrics. These metrics provide visibility into the performance, activity, and resource utilization of the Fargate-backed cluster, helping developers optimize their workloads and reduce costs. By leveraging these metrics and tools, developers can gain a deeper understanding of their Fargate clusters and improve their overall efficiency and reliability.
Mar 01, 2021
3,576 words in the original blog post.
Amazon ECS and EKS clusters running on Fargate can be monitored using Amazon CloudWatch, which aggregates monitoring data from many AWS services. CloudWatch collects metrics from ECS on Fargate in two separate namespaces: `AWS/ECS` for resource reservation and utilization metrics, and `ECS/ContainerInsights` for custom metrics that describe the status of ECS tasks and the number of running services, containers, and deployments. Container Insights is a separate namespace from `ECS/ContainerInsights` and stores EKS performance and resource metrics only from EC2-backed EKS clusters, not Fargate-backed EKS clusters. CloudWatch Logs collects logs from AWS services, including ECS on Fargate, and provides a console for exploring logs and a query language called CloudWatch Logs Insights for querying and analyzing logs. The Kubernetes Metrics API can be used to view metrics generated by cAdvisor, which runs as part of the kubelet on each compute resource. Prometheus can also collect cluster and resource metrics from Kubernetes, including EKS clusters running on Fargate, and stores metrics in a timeseries database. Fluent Bit can be used to route ECS logs to CloudWatch Logs, and kubectl can be used to view logs emitted by containerized applications. Datadog is a platform that collects, analyzes, and alerts on comprehensive monitoring data from Fargate and other technologies in the stack.
Mar 01, 2021
3,407 words in the original blog post.