January 2021 Summaries
10 posts from Datadog
Filter
Month:
Year:
Post Summaries
Back to Blog
Microsoft's Azure IoT Edge service enables running containerized workloads on IoT devices, allowing organizations in various industries to process data at the edge of their cloud networks. Datadog now integrates with IoT Edge, providing a comprehensive monitoring solution for Azure users. The integration allows deploying the Datadog Agent across all connected IoT Edge devices and visualizing key metrics from devices, modules, and IoT Edge Hubs. Monitoring IoT Edge Agents and IoT Edge Hubs helps identify issues such as insufficient memory, lack of communication between devices and Azure IoT Hub, increased latency, and more. Datadog's integration with over 650 technologies enables users to surface correlations across their entire stack for troubleshooting purposes.
Jan 21, 2021
665 words in the original blog post.
Azure IoT Edge is a Microsoft Azure service that allows organizations to run containerized workloads on IoT devices. This service, combined with Azure IoT Hub and Datadog's integration, provides enhanced monitoring capabilities for IoT infrastructure. With this setup, users can track device and module metrics, logs, and performance in real-time, enabling them to quickly identify issues such as insufficient memory or communication failures between devices and the cloud. The integration also provides customizable dashboards and alerts, allowing teams to take immediate action to address problems before further issues arise. By leveraging Azure IoT Edge with Datadog's integration, users can gain unparalleled visibility into their entire IoT infrastructure, streamlining monitoring and troubleshooting efforts.
Jan 21, 2021
673 words in the original blog post.
Oracle's Container Engine for Kubernetes (OKE) is a managed service that enables organizations to deploy, manage, and scale Kubernetes clusters in the cloud. By partnering with Datadog, users can gain comprehensive visibility into their OKE container infrastructure, monitor live processes, and track key metrics from all pods and containers in one place. The integration allows for monitoring of Kubernetes clusters on Oracle Cloud Infrastructure, tracking load on clusters, pods, and individual nodes, and automatic collection and reporting of metrics from services running in the cluster. Datadog also provides built-in Kubernetes dashboards, Live Container view for real-time insights into containers' resource consumption, Live Process view to track processes within a container, and Autodiscovery feature to monitor containerized services automatically.
Jan 20, 2021
925 words in the original blog post.
Istio is an open source service mesh that provides network abstraction for applications, enabling features like canary deployments and circuit breakers. Datadog's Network Performance Monitoring (NPM) helps visualize the topology of Istio-managed networks, monitor their health and performance, and identify issues related to Istio networking. NPM automatically accounts for Istio's network address translation logic, providing complete visibility into Istio traffic without any configuration. It also allows users to monitor key network performance metrics at the container, pod, and service layers, and correlate them with data from applications and infrastructure running in the mesh.
Jan 15, 2021
1,415 words in the original blog post.
Datadog Network Performance Monitoring is an automated tool that provides instant visibility into the topology of your Istio-managed network, allowing you to identify dependencies between services, pods, and containers. It helps dev and ops teams ensure that their Istio network infrastructure is healthy, performant, and routing traffic as intended. The tool automatically visualizes network performance metrics, such as volume, errors, and latency, and provides context on issues related to Istio's control plane and applications running in the mesh. By leveraging Datadog Network Performance Monitoring, teams can root out inefficiencies and misconfigurations, identify network errors and latency, pinpoint Envoy-related performance degradations, and get the context they need to locate issues between the Istio control plane and applications it manages.
Jan 15, 2021
1,448 words in the original blog post.
Serverless platforms like AWS Lambda have revolutionized application development by eliminating the need to manage infrastructure resources. However, serverless architecture presents new challenges in monitoring and troubleshooting applications. Datadog has introduced automatically-generated insights to provide deeper visibility into the health and performance of your functions. These insights help identify issues such as high memory usage, cold starts, or over-provisioned memory, enabling developers to quickly resolve problems and optimize resource allocation. Additionally, Datadog now displays error flags for Lambda functions, allowing users to drill down into individual function invocations for more granular troubleshooting. By tying functions to their associated traces and logs, Datadog makes it easier to gain context around each function's invocation and improve overall application performance.
Jan 14, 2021
772 words in the original blog post.
NVIDIA Jetson is a family of embedded computing boards designed for machine learning and AI applications at the edge. These devices are used in various sectors such as video and image processing, automating build processes in factories, and improving city infrastructures. Datadog IoT Agent now supports the current portfolio of Jetson boards, providing more visibility into IoT environments. It captures critical performance metrics from Jetson hardware, including GPU utilization and frequency, memory dedicated to the GPU, and external memory controller utilization. Datadog also collects standard system metrics for CPU, memory, and network I/O. This helps organizations monitor their fleet of Jetson devices and maintain them effectively. The platform provides full visibility into IoT networks, enabling quick identification of issues such as poorly performing or offline devices. It also allows users to proactively monitor their network with alerts that notify them when a device goes offline or experiences unusual drops in GPU utilization. Datadog helps organizations monitor critical resource metrics for their devices and track how specific events like software updates might have affected key device metrics, ensuring optimal performance of the fleet.
Jan 14, 2021
811 words in the original blog post.
Datadog has introduced automatically-generated insights to provide deeper visibility into the health and performance of AWS Lambda functions, enabling developers to quickly identify issues, troubleshoot errors, and optimize resource allocation. These insights use key data from Lambda functions to flag those that are failing or performing poorly, providing context into the nature of the problem, such as high memory usage or cold starts. The insights also include additional UI features, allowing developers to pivot to relevant traces and logs for immediate troubleshooting. Additionally, Datadog ties functions to their associated traces and logs, making it easier to correlate failing Lambda functions with relevant data, spot invocation problems in real-time, and view serverless functions in full view alongside other technologies.
Jan 14, 2021
789 words in the original blog post.
Datadog has now added support for NVIDIA Jetson boards to its IoT Agent, providing visibility into critical performance metrics such as GPU utilization and memory allocation. This allows users to monitor their fleet of Jetson devices in real-time, gaining insights into how they are performing and identifying potential issues before they become problems. With this integration, organizations can optimize their device networks, proactively monitor for performance issues, and ensure that their resource-intensive workflows are running smoothly. The addition of Datadog's IoT Agent support for Jetson boards provides end-to-end visibility into the entire network, enabling users to make informed decisions about their devices and infrastructure.
Jan 14, 2021
822 words in the original blog post.
Watchdog has made its Root Cause Analysis (RCA) feature available in general availability, allowing users to automatically identify the root cause of anomalies and critical failures in their applications and infrastructure. This hands-free approach enables users to resolve problems faster than ever, reducing mean time to resolution (MTTR). Watchdog RCA maps applications and infrastructure, understanding how they interact, and uses this knowledge to identify probable root causes and surface resulting critical failures. The feature can detect a range of root causes, including problematic code changes, increased traffic from clients, AWS instance failures, and disk capacity issues. Additionally, Watchdog Impact Analysis assesses the end-user impact of issues, allowing users to prioritize troubleshooting efforts based on which services are affecting the most users. By automating root cause analysis, users can resolve issues quickly, minimizing their effect on the end-user experience.
Jan 05, 2021
809 words in the original blog post.