December 2025 Summaries
41 posts from Datadog
Filter
Month:
Year:
Post Summaries
Back to Blog
Security teams face challenges due to the diverse and inconsistent log formats across various systems they defend, hindering their ability to efficiently correlate events and investigate incidents. Datadog Cloud SIEM addresses this by leveraging the Open Cybersecurity Schema Framework (OCSF) for normalizing security telemetry, recently expanding its coverage to include more sources such as AWS CloudTrail, Okta, and GitHub. The new OCSF processor offers self-service control, allowing teams to map any log source to OCSF without custom engineering, thus standardizing events for consistent analysis. This facilitates the creation of unified detection rules across multiple data sources, reducing operational overhead and enabling faster and clearer incident triage. The processor supports an open SIEM operating model, giving teams flexibility in data collection and routing, and reducing schema lock-in. By extending normalization to any log source, Datadog allows security teams to integrate new data sources seamlessly and manage growing environments with reduced complexity.
Dec 30, 2025
701 words in the original blog post.
Kevin Francis, a Platform Engineering Manager at Cambia Health Solutions, describes how his team successfully merged two internal divisions to optimize cloud costs using Datadog’s Cloud Cost Management (CCM) and Resource Catalog. The merger aimed to unify observability, which was fragmented across different platforms, hindering cost reduction and operational awareness. By standardizing on Datadog, the team achieved a comprehensive view of their infrastructure, identifying nearly $30,000 in monthly savings through a refined Reserved Instance (RI) purchasing strategy for Amazon Relational Database Service (RDS). This process involved standardizing the RDS fleet on Graviton-powered db.r6g instances and centralizing RI purchasing, which allowed for greater discount sharing and size flexibility, thereby maximizing cost savings. The initiative also utilized Datadog dashboards for monitoring savings and Resource Catalog for preventing configuration drift, ultimately establishing a repeatable framework for future optimizations in other resources like DynamoDB and OpenSearch.
Dec 30, 2025
1,075 words in the original blog post.
Engineers at Zendesk, including Anatoly Mikhaylov and Nick Hefty, detailed their approach to optimizing the cost of observability data while maintaining essential visibility for troubleshooting. By leveraging Datadog's tools, they centralized metrics and monitors and adopted Application Performance Monitoring (APM) and Log Management for deeper insights, leading to enhanced incident response and performance. The team conducted an observability audit using the Pareto Principle to identify data that was both valuable and costly, focusing on optimizing trace and log usage. They implemented single-span ingestion and custom facets to enrich root spans, reducing data volume without losing critical insights. They also experimented with log reduction strategies, such as batching log entries and applying exclusion filters, to cut costs. These efforts led to a significant reduction in observability expenses and a clearer understanding of cost distribution across services, without disrupting engineering workflows. The initiative successfully demonstrated that strategic optimizations can sustain critical operational data access while controlling financial impact.
Dec 26, 2025
3,804 words in the original blog post.
Datadog offers a comprehensive platform to help organizations manage the costs, performance, and infrastructure efficiency of AI applications as they transition to production-scale workloads. By integrating Cloud Cost Management, LLM Observability, and GPU Monitoring, Datadog provides real-time visibility across the AI stack, allowing organizations to connect spending to performance and ensure resources are used effectively. This platform enables finance, engineering, and operations leaders to share a unified view of AI expenditures and make informed decisions in real time. Datadog's LLM Observability helps teams evaluate and enhance AI application performance, ensuring financial efficiency without compromising quality. Additionally, GPU Monitoring offers insights into GPU usage, helping to reduce waste and improve efficiency. Datadog's unified approach links application and infrastructure metrics, facilitating real-time decisions that optimize AI spending and performance, ultimately increasing AI ROI.
Dec 23, 2025
1,146 words in the original blog post.
Datadog's Cloud Network Monitoring (CNM) Network Health is a tool designed to provide a unified view of network issues across various specialized domains within organizations, helping to reduce mean time to resolution (MTTR) by clearly identifying the source of connectivity problems and providing contextual recommendations for resolution. The tool enhances visibility by correlating real-time flow data, TCP metrics, and cloud configurations to pinpoint issues such as security group misconfigurations, TLS handshake errors, and DNS resolution failures, which traditional observability tools may overlook. By leveraging insights from the Datadog Watchdog AI engine, Network Health offers in-depth analysis and suggested remediation steps, enabling teams to address the root causes of network problems more efficiently and ensuring that network connections remain secure and reliable.
Dec 23, 2025
777 words in the original blog post.
Account takeovers (ATOs) pose a significant threat to online platforms, with attackers often exploiting leaked credentials through techniques like credential stuffing. Datadog has developed an automated system to combat ATOs by conducting real-time checks during login events to identify compromised credentials, using a k-anonymity technique to ensure user privacy. This approach allows Datadog to reset passwords proactively when a breach is detected, reducing the need for manual intervention and minimizing disruption to the user experience. The system, which focuses on password-based authentication, operates in the background to provide enhanced security without impacting login times. Future plans include expanding protections beyond passwords to address threats like cookie theft and deepening integration with the Safety Center for better visibility into credential exposures.
Dec 22, 2025
1,100 words in the original blog post.
In the late 2000s, the debate over SQL versus NoSQL databases was prevalent, with predictions of a winner-take-all outcome. However, the rise of microservice architectures has led organizations to adopt multiple database technologies, allowing teams to choose solutions tailored to their specific needs. This shift from monolithic to microservice architectures has fragmented databases across services, enabling organizations to use a "right tool for the job" approach and adopt a mix of SQL and NoSQL databases. While this flexibility offers development advantages, it also introduces challenges such as schema fragmentation, complex service integration, and the need for data integration layers like GraphQL. Organizations are increasingly using message queues to decouple services and enhance system resiliency, while cloud-based data platforms like Snowflake and Redshift facilitate analytics by consolidating data from distributed systems. The transition reflects a broader trend of using diverse database technologies to optimize performance and cost-effectiveness in handling various use cases.
Dec 22, 2025
2,133 words in the original blog post.
Early-stage engineering teams often prioritize speed, which, while advantageous, can result in numerous disruptive signals like stack traces and timeouts. Differentiating between these signals and identifying those that critically impact users and revenue is crucial. By integrating engineering telemetry data, such as logs and user sessions, with customer feedback from support tickets and reports, teams can form a cohesive feedback loop. This approach allows them to evaluate defects based on who was affected, where in the user journey the issue occurred, and its impact on conversion, engagement, or revenue. Implementing a feedback loop with tools like Datadog enables teams to prioritize issues based on their real-world impact, turning incidents into valuable product insights. Standardizing correlation identifiers like trace_id and user.id across telemetry data ensures seamless analysis and error tracking, while structured logging and customer feedback enhance the visibility and prioritization of impactful defects. By focusing on how errors affect core product flows, engineering and product teams can collaborate more effectively, ensuring improvements are based on evidence-backed insights rather than visibility alone, ultimately enhancing user experience and business outcomes.
Dec 19, 2025
3,816 words in the original blog post.
Dieter Matzion, a seasoned cloud practitioner, discusses the importance of managing custom metrics in Datadog to ensure they provide true business value while staying within budgetary constraints. As organizations grow, proactive governance and cost awareness become crucial, and aligning finance with engineering is essential for effective cloud cost management. Matzion emphasizes the role of unit economics in linking custom metrics usage to revenue-driving or cost-saving activities, creating a shared language between technical and financial stakeholders. He highlights how Datadog's tools, such as standardized metric-naming conventions and the Metrics without Limits feature, can help reduce unnecessary volumes and optimize custom metrics by excluding unused tags. By doing so, organizations can enhance observability and maintain efficient monitoring practices. Matzion also notes the importance of tracking overall spend trends to catch abnormalities early and suggests holding regular finance reviews to address budget issues.
Dec 19, 2025
898 words in the original blog post.
APIs are central to modern digital products but pose significant security challenges due to their expanding endpoints and increased attack surfaces. Datadog's App and API Protection (AAP) addresses these issues by providing comprehensive API visibility and security solutions, integrating discovery, monitoring, and remediation within a single platform. AAP features an API Inventory that consolidates data from multiple sources, including live traffic and source code analysis, to offer a unified view of an organization's API landscape. It detects misconfigurations and vulnerabilities through API Findings, which prioritize risks and automate responses. The API posture overview offers an organization-wide assessment of security health, highlighting undocumented and high-risk endpoints, while built-in detection rules and custom queries enable continuous monitoring and protection. AAP's integration with Datadog's existing tools allows for effective alerting and remediation, turning API visibility into proactive security management and facilitating the evolution of API practices to align with modern risk models.
Dec 18, 2025
1,450 words in the original blog post.
Datadog's Fleet Automation streamlines the setup and scaling of observability across distributed environments by allowing platform and SRE teams to manage Datadog products and Agent integrations centrally, reducing the need for manual configuration and coordination. This solution enables teams to configure monitoring from a single platform, ensuring complete observability coverage by identifying and addressing telemetry gaps and misconfigurations. Fleet Automation offers features like guided workflows, code-driven configuration, and an API for integration with existing automation tools, simplifying the deployment and management of configurations. It also provides role-based access control, audit trails, and deployment controls to handle complex environments effectively, allowing teams to maintain visibility and consistency across their infrastructure while minimizing onboarding time and management overhead. The centralized interface aids in troubleshooting by offering insights into telemetry data reporting and configuration history, supporting distributed teams with notifications and detailed event logs.
Dec 18, 2025
821 words in the original blog post.
Operating Cilium at scale involves meticulous configuration and monitoring to maintain reliability across hundreds of Kubernetes clusters, thousands of nodes, and pods in multi-cloud environments, as demonstrated by Datadog's experience. Key to this process is the adoption of native routing to minimize overhead, the standardization of operator/agent splits to manage cloud API interactions, and fine-tuning IP Address Management (IPAM) to optimize resource allocation. Additionally, stable identity management, consistent Maximum Transmission Unit (MTU) settings, and rigorous upgrade validation practices are crucial for preventing disruptions. Datadog emphasizes the importance of monitoring Cilium's control plane and datapath signals to preemptively address issues and ensure operational stability. By leveraging tools like bpftrace and bpftool, they investigate and rectify datapath anomalies, and they adopt the kube-proxy replacement to enhance service load balancing via eBPF. Overall, Datadog's approach highlights the significance of standardized practices in managing large-scale deployments, ensuring that even the largest clusters operate smoothly and predictably.
Dec 18, 2025
3,671 words in the original blog post.
Continuous profiling has become a crucial aspect of observability, often referred to as the fourth pillar, yet it can be challenging for newcomers to implement effectively, especially in distinguishing between memory allocation and retention. The blog post explains how flame graphs can sometimes mislead developers into incorrectly identifying memory issues, as they highlight where memory is allocated rather than where it is retained. It offers insights into using Python code to discern between memory allocation versus retention, emphasizing that this differentiation is important for accurate troubleshooting. Through examples like the `allocator_vs_holder.py` and the functions `grow_heap` and `heavy_churn`, the post illustrates how to identify memory retainers and choose between heap live size and allocated memory views based on the nature of the issue. It also highlights the utility of the Datadog Continuous Profiler in visualizing memory usage over time, aiding in the identification of memory retaining paths, which can differ from allocation paths. The blog notes that understanding these paths requires a deep knowledge of the source code, and mentions that Datadog is developing solutions to identify memory retaining paths for Java and .NET, encouraging users to explore their profiling tools and try a 14-day free trial.
Dec 18, 2025
1,576 words in the original blog post.
Cilium network policies (CNPs) enhance Kubernetes by extending L3/L4 controls to the application layer (L7), offering advanced networking capabilities that can also introduce new connectivity challenges, especially in large environments. These challenges often arise from differences in how Kubernetes and Cilium interpret concepts like label scoping and IP-based rules, impacting areas such as cross-cluster communication, egress rules, CIDR-based L3 policies, and namespace isolation. Misconfigurations can lead to unintended traffic blocking or allowance, requiring careful policy structuring and understanding of Cilium's unique handling of cluster entities and security identities. To address these issues, Cilium users must navigate policy settings such as the policy-default-local-cluster and use appropriate selectors like endpoints-based policies for cross-cluster traffic or entity-based policies for internal traffic, ensuring correct implementation for desired connectivity outcomes. Awareness of these differences and proper monitoring using tools like Hubble can help in diagnosing and resolving common misconfigurations, ultimately leading to more reliable and secure network policies.
Dec 16, 2025
2,054 words in the original blog post.
Connecting development work to real user outcomes is often challenging for engineering and product teams due to fragmented monitoring systems, which can hinder effective project prioritization. Both teams aim to improve user experience but use different metrics and tools, leading to a disconnect in understanding app performance and user engagement. Frontend teams focus on system performance indicators like Core Web Vitals and error rates, while product teams prioritize user engagement metrics such as conversion and retention rates. A shared context, achieved by integrating Real User Monitoring (RUM) and product analytics, can bridge this gap by unifying performance and engagement data. Using common datasets and tools like Datadog, teams can synchronize their goals and collaborate more effectively by aligning on shared user action definitions. This approach aids in troubleshooting, decision-making, and prioritizing projects that enhance performance and user engagement. Enhanced collaboration allows engineers to see the impact of their work, while product teams better understand technical challenges, fostering a holistic view of how releases influence user journeys.
Dec 15, 2025
1,112 words in the original blog post.
Kubernetes operators are crucial in managing application behavior by automating tasks such as scaling, upgrading, and failure recovery, with their performance directly affecting the applications they oversee. These operators are built on the Kubernetes controller pattern, where they continuously reconcile the application's current state with its desired state using a reconciliation loop, thus ensuring the application's consistency. Monitoring the performance of these operators through metrics and logs is vital for maintaining application reliability, as it helps in identifying issues like latency, errors, and resource exhaustion. Metrics such as reconciliation attempts, loop duration, and work queue depth provide insights into operator efficiency, while Go runtime metrics can reveal inefficiencies like memory leaks. Tools like Prometheus and Datadog are used to collect, store, and visualize these metrics, allowing for proactive diagnosis and resolution of performance issues. By leveraging these monitoring capabilities, organizations can ensure their Kubernetes operators remain reliable and efficient, ultimately supporting the smooth operation of the applications they manage.
Dec 15, 2025
2,667 words in the original blog post.
In 2025, the cloud security landscape faced enduring concerns alongside emerging challenges, primarily driven by increased AI adoption and evolving attacker strategies. The rapid integration of AI technologies introduced new vulnerabilities due to unpredictable user input and nascent security models, while persistent issues like long-lived credentials and third-party package vulnerabilities continued to threaten cloud environments. Attackers increasingly targeted identities, development pipelines, and AI tools, exploiting security gaps within shifting cloud perimeters defined by data rather than networks. Notably, there was a marked focus on supply chain attacks, particularly within developer environments and CI/CD pipelines, exemplified by incidents like the npm worm attack. Organizations were encouraged to enhance their monitoring systems, integrate security with incident management, and fortify defenses by minimizing credential lifespans, securing AI systems, and protecting supply chain components to mitigate these evolving threats.
Dec 15, 2025
1,175 words in the original blog post.
AWS re:Invent 2025 highlighted significant advancements in cloud technology, emphasizing three critical domains: applied intelligence, trust and governance, and enterprise velocity. The event showcased innovations in agentic AI, including the launch of AWS Trainium 3 and Amazon Bedrock AgentCore, which facilitate faster model training and deployment, although accountability and security remain crucial challenges. AWS introduced enhanced observability and security features like LLM Observability and Bedrock AgentCore Policies to ensure safe scaling, while IAM Policy Autopilot and AWS Security Agent aimed to streamline security operations. To boost enterprise velocity, new AI capabilities were unveiled for Code Security, alongside tools like Amazon ECS Express Mode and AWS Lambda Managed Instances, designed to simplify and accelerate deployment processes. The event underlined the importance of integrating AI, security, and developer velocity into a unified operational strategy, while hands-on workshops and sessions enriched attendee experiences and highlighted practical applications of these technologies.
Dec 11, 2025
1,818 words in the original blog post.
Datadog's latest episode, focused on releases announced at AWS re:Invent 2025, highlights several innovations aimed at enhancing cloud management and security. Key features include the MCP Server, which connects AI agents like Amazon Kiro to Datadog's tools for real-time data analysis, and Bits AI SRE, which streamlines alert investigation and incident coordination. CloudPrem, a hybrid log management solution, offers complete visibility and retention control within user infrastructures, while Storage Management provides detailed cost insights for Amazon S3 buckets. Agent Builder enables the creation of AI agents for complex decision-making, and Secret Scanning helps detect credential leaks. Additional updates cover governance, AI agent behavior visibility, Kubernetes remediation, and LLM observability, with a promise of ongoing feature releases and updates in future episodes.
Dec 11, 2025
718 words in the original blog post.
Datadog Teams' integration with GitHub allows engineering organizations to import and sync team and membership data to maintain accurate service ownership information, which is crucial for effective operations. By connecting GitHub team structures to the Datadog Internal Developer Portal, companies can automate the updating of team hierarchies and service ownership, reducing manual efforts and the risk of outdated information. This integration offers a centralized view of team data, enhancing accountability and operational efficiency by enabling developers, platform teams, and engineering leaders to access relevant metrics, ownership details, and performance reports. The system supports better collaboration, incident management, and adherence to operational standards by providing a continuously updated map of team structures and responsibilities, ultimately improving onboarding processes and the overall health of engineering workflows.
Dec 09, 2025
800 words in the original blog post.
Success for security organizations, unlike other engineering teams that measure success through tangible outcomes, is defined by reducing risks and improving the company's security posture over time. As companies grow and adopt new technologies, such as AI and edge environments, their attack surfaces expand, requiring security to adapt accordingly. Datadog's approach involved integrating its SRE and security groups, emphasizing the creation of scalable, secure systems, maintaining connected specialized teams, and preparing for emerging technologies. Leadership at various levels plays a critical role in defining priorities, implementing security guardrails, and ensuring effective communication and alignment across teams. This includes embedding AI responsibly, setting clear ownership, and developing governance guidelines. Datadog measures success through metrics like detection signal accuracy, ensuring that systems remain resilient and responsive to new risks while supporting broader company goals.
Dec 08, 2025
1,346 words in the original blog post.
Datadog Infrastructure Management in Preview addresses the challenges organizations face in managing cloud infrastructure changes by offering proactive tools for detecting and remediating configuration issues. As modern environments grow increasingly complex, traditional methods like audits and notifications often fall short in maintaining well-architected standards, leading to increased risks and costs. Datadog's solution integrates change detection, policy evaluation, and automated remediation, allowing infrastructure teams to efficiently manage, assess, and enforce best practices across various cloud platforms such as AWS, Azure, Google Cloud, and Kubernetes. It prioritizes risky changes using historical data and configuration patterns, enabling teams to focus on significant issues while minimizing noise. By centralizing policy creation and assessment, it helps maintain consistent infrastructure standards and reduces operational overhead. Additionally, Datadog facilitates automated issue resolution through integration with tools like Jira and ServiceNow, and by connecting resources to their infrastructure-as-code definitions, which accelerates the remediation process.
Dec 04, 2025
910 words in the original blog post.
Datadog, originally a single-product company focused on bridging the gap between development and operations, has evolved its observability platform to address the challenges and opportunities presented by the rapid adoption of AI. As AI introduces new risks and monitoring challenges, Datadog is developing AI-powered features and tailored monitoring tools to help organizations improve their AI systems. The company invests significantly in R&D to create agentic AI solutions, like Bits AI agents, which autonomously manage incidents and security within complex environments. Datadog's observability suite now includes advanced tools like Model Context Protocol, Datadog MCP Server, and Toto, a timeseries foundational model, to enhance AI, ML, anomaly detection, and forecasting capabilities. These efforts aim to support organizations in navigating the complexities of AI at scale, enabling them to innovate rapidly while maintaining performance, security, and compliance. By engaging with both startups and major AI players, Datadog aims to remain at the forefront of AI innovations, ultimately striving for a future of proactive operations and security management that could lead to zero-incident environments.
Dec 02, 2025
1,107 words in the original blog post.
Datadog's Kubernetes Cluster Autoscaler, currently in limited preview, addresses the challenge of overprovisioning and underutilization in Kubernetes infrastructures by offering automated solutions to optimize compute resources and reduce cloud costs. By simulating existing clusters, it provides recommendations for cost-efficient node configurations and enables automatic workload migration through managed integrations like Karpenter or GitOps solutions. The tool leverages Datadog's Kubernetes observability to align node capacity with real workload behavior, ensuring performance and availability are not compromised. It also considers constraints such as node and pod affinities, taints, and disruption budgets, offering insights into potential savings and workload impacts before changes are made. For GPU-backed AI and ML workloads, it provides specific recommendations to match actual usage patterns. The system supports both live node scaling and GitOps workflows, allowing organizations to update node configurations efficiently while maintaining application performance.
Dec 02, 2025
881 words in the original blog post.
As development teams rapidly integrate generative AI, security teams encounter new challenges in safeguarding the software development life cycle, particularly as legacy scanning tools struggle to keep pace with the increasing speed and scale of code changes. Datadog Code Security addresses these challenges by utilizing AI-driven automation to combine static and runtime analysis, effectively scanning repositories for vulnerabilities in first-party code, open-source dependencies, and infrastructure-as-code misconfigurations. It excels in detecting hidden code vulnerabilities and validating findings by filtering out false positives to reduce alert fatigue and improve remediation time. By employing large language models (LLMs), Code Security evaluates code behavior in context, identifying risky code changes in pull requests that traditional static analyzers might miss. The system allows for the prioritization of high-risk findings, enabling teams to focus on the most actionable issues by providing transparency in vulnerability classifications. Moreover, Code Security facilitates batch remediation by generating proposed code patches through AI collaboration, allowing developers to efficiently resolve vulnerabilities without disrupting their workflow. This modern approach integrates AI-native analysis with a focus on developer experience, offering a comprehensive solution to secure applications in today's fast-paced development environment.
Dec 01, 2025
840 words in the original blog post.
As generative AI workloads increase, engineering and platform teams are adopting OpenTelemetry (OTel) to standardize observability by providing a unified pipeline for telemetry data across applications, infrastructure, and AI systems. OpenTelemetry GenAI Semantic Conventions offer a standardized schema to track and analyze AI workloads, making them measurable and interoperable across frameworks. Datadog has integrated native support for these conventions, allowing seamless instrumentation of LLM applications that can be analyzed within Datadog LLM Observability without additional code modifications. This integration enables teams to forward GenAI spans directly to Datadog, ensuring that data governance policies are maintained while providing comprehensive visibility into AI performance, quality, and cost metrics across various providers and models. By mapping GenAI attributes to Datadog's schema, teams can analyze token usage, latency, and cost, correlating AI data with broader application performance metrics. The platform also supports experimentation with agentic applications, offering tools to iterate and refine AI systems efficiently.
Dec 01, 2025
939 words in the original blog post.
The Model Context Protocol (MCP) is crucial for connecting AI agents to external tools and data, and understanding the behavior of MCP clients, such as agents, gateways, and IDEs, is vital for efficient system operation. Datadog's LLM Observability now offers comprehensive tracing and monitoring for these clients, capturing every step from session initialization to tool invocation as part of a span linked to the LLM trace that initiated the tool selection. This enhanced visibility helps teams trace client-side registry discovery, tool invocation behavior, and their contribution to latency and token usage, thereby pinpointing failures and measuring efficiency. By automatically instrumenting the MCP Python client library, teams gain insights into the complete MCP life cycle, enabling them to identify slow or unreliable MCP servers, track connection latency, and correlate failures with specific tools or prompts. Additionally, the observability platform aggregates MCP span data to provide key performance metrics, such as latency and error rates, allowing engineers to improve registry configurations and enhance AI agent performance. This level of detail supports quicker issue resolution, reduced unnecessary tool invocations, and optimized MCP integrations in production environments.
Dec 01, 2025
964 words in the original blog post.
AI coding assistants, like Claude Code, are becoming essential tools in software engineering, enhancing developers' ability to write, refactor, and review code more efficiently. However, monitoring their performance and utility is crucial for ensuring they provide value. Datadog's new Claude Code Monitoring, available in the AI Agents Console, offers organizations a comprehensive view of AI agent and coding assistant usage, including performance, reliability, and cost analysis. This feature allows teams to track real-time telemetry data, analyze error rates and latency patterns, and evaluate spend and ROI efficiently. By providing detailed insights into user activity, session durations, and code change activity, the console helps platform and engineering leaders make informed decisions on AI adoption, ensuring it aligns with organizational goals and budgets. The monitoring tool also allows for the identification of anomalies in usage and spend, helping to optimize configurations for cost-efficiency without compromising developer experience.
Dec 01, 2025
1,077 words in the original blog post.
Amazon Elastic Container Service (ECS) Managed Instances offer a fully managed compute option for running containerized workloads on Amazon EC2, alleviating the operational overhead of infrastructure management by handling provisioning, patching, scaling, and maintenance. This service provides developers with access to a wide variety of EC2 instance types, including those with GPU acceleration and high-throughput networking capabilities, which are particularly useful for AI and data processing workloads. Datadog has expanded its ECS monitoring capabilities to fully support ECS Managed Instances, offering tools to analyze cluster performance, troubleshoot issues, and correlate telemetry data across ECS environments. The ECS Explorer in Datadog provides a unified view of ECS resources, allowing users to inspect resource configurations and analyze performance signals, while the OOTB Amazon ECS dashboard offers a high-level overview of the environment. Datadog's monitoring features include alert configurations for ECS tasks and the ability to quickly navigate between related resources, thereby improving visibility and reliability across ECS deployments.
Dec 01, 2025
888 words in the original blog post.
Prompt Tracking, a feature available in Datadog LLM Observability, provides a systematic way to manage and observe changes in prompts used by LLM agents and applications, which are crucial for refining performance and accuracy. This feature allows prompts to be defined, versioned, and monitored as first-class artifacts, enabling teams to apply rigorous management akin to that of application code and models. Prompt Tracking offers the ability to correlate prompt changes directly with performance metrics such as error rate, token usage, and latency, thus facilitating informed rollout decisions based on real data. By integrating with tools like the LangChain framework, it automatically instruments prompt IDs and versions, providing an auditable link between prompt iterations and their performance impacts. The feature also includes a Prompt Playground for validating changes before production, and a consolidated analytics dashboard for tracking prompt versions and trends. This structured approach aligns prompt development with established software practices, enhancing visibility and reducing the risk of regressions in LLM applications.
Dec 01, 2025
777 words in the original blog post.
Datadog AI Research is participating as a Gold-level sponsor at NeurIPS 2025, showcasing its work on observability-native foundation models and AI agents. They are presenting Toto, a time series foundation model, and BOOM, a large-scale benchmark, at a poster session on December 5 in the San Diego Convention Center. Datadog is also involved in the BERT²S workshop on December 7, which discusses benchmarking and the real-world impact of time series models. Interested attendees can visit Booth 631 in the Expo Hall from December 2-5 to learn about Datadog's AI projects, including autonomous SRE agents and AI-powered developer tools. The team is led by Ameet Talwalkar and Chenghao Liu, who focus on projects in cloud observability and security, inviting potential collaborators and job seekers to explore roles in AI research.
Dec 01, 2025
372 words in the original blog post.
Datadog Kubernetes Autoscaling introduces a flexible approach to scaling workloads by allowing users to utilize any integration or custom metrics collected in Datadog, rather than relying solely on CPU or memory metrics. This approach is particularly beneficial for applications, such as those downstream of messaging systems or data pipelines, where traditional resource utilization metrics do not accurately reflect workload demands. Custom Query Scaling enables users to define precise scaling policies using Datadog's query editor, facilitating the use of application-specific metrics like message throughput or request rates. This method not only enhances responsiveness and resource allocation but also simplifies cluster management by eliminating the need for additional components, centralizing scaling configuration and performance monitoring within the Datadog platform. By aligning scaling with actual application performance, Datadog Kubernetes Autoscaling improves both reliability and cost-efficiency, offering a streamlined solution for managing Kubernetes workloads.
Dec 01, 2025
699 words in the original blog post.
Serverless computing has transformed application development by eliminating infrastructure management, allowing developers to concentrate on coding. AWS Lambda, a prominent player in this realm, executes code in response to events, automatically scaling and managing infrastructure, and charging based on usage, making it suitable for varying traffic patterns. However, for intensive or steady workloads, many opt for Amazon EC2 due to its consistent performance and hardware flexibility. To address this, AWS introduced Lambda Managed Instances, combining Lambda's serverless model with EC2's hardware versatility, which is especially beneficial for tasks requiring specific resources like GPUs. Datadog, an AWS partner, provides comprehensive observability for these instances, unifying metrics, logs, and traces to aid in identifying performance bottlenecks and optimizing workloads. Organizations can monitor Lambda Managed Instances alongside other services using Datadog's dashboards, which facilitate a seamless transition for functions moving from Lambda's traditional model to this new infrastructure. This integration allows for detailed insights into performance, resource usage, and cost management, ensuring teams can efficiently manage and optimize their serverless applications.
Dec 01, 2025
774 words in the original blog post.
Datadog has introduced an AI-powered log parsing feature in its Log Explorer to streamline the process of analyzing complex, unstructured log data. This tool automates the generation of Grok parsing rules, which extract key information from raw log lines into structured fields, enabling deep analysis without altering global ingestion pipelines or relying on manual regex testing. In distributed microservices environments, where a single user request can generate numerous log events across multiple systems, this feature reduces the time and effort traditionally needed to interpret logs, craft parsing patterns, and verify their accuracy. The AI-driven approach allows users to auto-extract fields, standardize logs from various sources, and perform advanced operations like filtering and geospatial analysis, thus facilitating faster insights and reducing the need for context switching during investigations. By transforming log standardization into a flexible, query-time activity, Datadog enhances the capability of engineers and analysts to efficiently debug, investigate incidents, and analyze network traffic.
Dec 01, 2025
670 words in the original blog post.
Datadog Observability Pipelines has introduced metric support, currently in Preview, offering organizations a unified platform to manage telemetry data governance for both logs and metrics, addressing challenges like separate tools, inconsistent tags, and high costs from low-value logs. This update allows teams to control ingestion costs by applying consistent tag policies and filtering out unnecessary or malformed metrics. The integration with existing tools, such as the Datadog Agent and OpenTelemetry Collector, ensures seamless governance without requiring new tools or workflows. It provides granular control over metrics, enabling the definition of rules for inclusion and exclusion based on names or tag patterns, thereby simplifying operational processes and improving data quality and consistency. Additionally, built-in diagnostics offer insights into pipeline health, allowing teams to monitor and verify pipeline operations effectively. This consolidated approach reduces tool sprawl, enhances data consistency, and provides a more reliable foundation for observability across various environments.
Dec 01, 2025
750 words in the original blog post.
OpenTelemetry (OTel) has established itself as the standard for telemetry data collection and routing, yet its large-scale adoption presents operational challenges due to distributed collector fleets that can suffer from configuration drift and version sprawl. Datadog Fleet Automation addresses these issues by providing centralized visibility and management of OTel collectors, enabling teams to monitor fleet status, inspect configurations, and quickly identify inconsistencies. This tool aggregates rich metadata for each collector, supports search and filtering functions to spot outdated versions, and simplifies configuration troubleshooting by integrating live configuration data directly into Datadog. By streamlining these processes, Fleet Automation reduces the operational overhead of managing OTel pipelines, enhances telemetry pipeline control, and accelerates issue resolution, thereby supporting organizations in confidently scaling their OTel deployments.
Dec 01, 2025
651 words in the original blog post.
The DDOT gateway, now in public preview, enhances the Datadog Distribution of the OpenTelemetry (OTel) Collector by providing a scalable gateway deployment model for centralized management of telemetry pipelines. This solution addresses challenges faced by organizations in maintaining consistency across extensive services and teams by enabling centralized governance, advanced data pre-processing, and multi-destination routing from a single source. Unlike node-level collectors, the gateway allows platform and SRE teams to enforce consistent tagging and schema standards, apply sophisticated processing like tail-based sampling, and manage routing logic efficiently. It provides elastic scalability through Kubernetes' Horizontal Pod Autoscaler, ensuring high throughput and preventing data loss during traffic surges. Additionally, the DDOT gateway offers full lifecycle management, allowing teams to monitor performance and resource usage seamlessly, supported by Datadog's Fleet Automation for managing configurations and detecting configuration drift. This centralized approach facilitates the creation of a unified and resilient OpenTelemetry pipeline, optimizing data quality and cost management while maintaining robust observability infrastructure across dynamic environments.
Dec 01, 2025
1,256 words in the original blog post.
Strands Agents, an open-source Python framework developed by AWS, simplifies the creation of production-ready AI agents by abstracting orchestration and allowing models to plan and execute tasks autonomously. However, these applications pose challenges in maintaining visibility and predictability of multi-agent workflows, often resulting in performance issues and debugging complexities. Datadog LLM Observability addresses these challenges by providing out-of-the-box visibility for Strands Agents, enabling developers to trace, measure, and evaluate agent workflows without custom instrumentation. It captures key operations within the agent lifecycle, offering end-to-end traceability and correlating model performance with infrastructure health. Developers can debug workflows, evaluate safety and quality, and experiment with models using real production data, identifying inefficiencies and optimizing multi-agent processes. Additionally, Datadog's evaluation tools help detect errors such as hallucinations and unsafe responses, allowing for rapid diagnosis and improvement of AI applications. By integrating seamlessly with Datadog's monitoring products, this observability tool enhances the reliability and performance of AI agent deployments.
Dec 01, 2025
848 words in the original blog post.
Securing sensitive information in code is a challenging task, often complicated by developers hardcoding credentials or using AI-generated code that includes live API keys, leading to inadvertent exposure of enterprise secrets. Datadog Secret Scanning, now generally available as part of Datadog Code Security, addresses this issue by detecting, validating, and blocking exposed credentials to prevent security and compliance risks. Secret Scanning continuously monitors source code, repositories, and CI/CD pipelines for credential leaks, integrating directly into developer workflows to ensure rapid remediation and shift-left security. It prioritizes real, active credentials through live third-party validation, reducing false positives and alert fatigue, while proactively enforcing pre-commit and pre-merge checks to block secrets from entering codebases. This tool also correlates exposure data with runtime signals, vulnerabilities, and misconfigurations to provide a comprehensive view of application security, helping teams to trace exposures, revoke keys, and implement preventive policies. As part of the broader Code Security suite, Secret Scanning complements other tools such as Software Composition Analysis, Static and Runtime Code Analysis, and Infrastructure-as-Code Security to offer a unified risk management solution throughout the software development lifecycle.
Dec 01, 2025
628 words in the original blog post.
Datadog Cloud SIEM is a security information and event management solution designed to help large organizations manage security threats and operations efficiently across complex cloud and SaaS environments. The platform centralizes insights, accelerates threat detection, and automates responses by leveraging AI tools like Bits AI Security Analyst, which autonomously investigates signals. By integrating the Open Cybersecurity Schema Framework (OCSF), Datadog normalizes logs from diverse sources, providing a consistent structure for analysis and enabling prebuilt detection rules that apply across multiple platforms. The solution also includes features like Risk Insights for prioritizing alerts based on severity and frequency, and Sequence Detections to identify coordinated attack patterns. Additionally, Datadog offers integrations and Content Packs for fast onboarding and broad coverage, supporting over 90 platforms. The platform aims to reduce operational overhead and improve security teams' efficiency, as demonstrated by MyFitnessPal's successful migration, which resulted in significant cost savings and enhanced search performance.
Dec 01, 2025
1,693 words in the original blog post.
Datadog Cloud Security offers a comprehensive solution for understanding and mitigating risks in complex cloud environments by providing contextual insights into attack paths and vulnerabilities through its Security Graph. This tool helps security teams visualize the network paths and IAM relationships that could expose resources to potential attacks, allowing for a more informed prioritization of remediation efforts based on real-world exposure. By mapping out the public reachability and lateral movement possibilities, the Security Graph highlights how vulnerabilities could be exploited and the subsequent actions an attacker might take, such as escalating privileges or accessing additional data. This enables teams to focus on genuine risks by filtering out low-impact findings and assessing the potential blast radius of a compromise. By linking vulnerabilities to network components, IAM entities, and dependent resources, the Security Graph facilitates a unified and interactive view of the cloud environment, helping engineers validate reachability and coordinate remediation efforts more efficiently. Currently in preview for AWS environments, this feature is available to Datadog Cloud Security customers, who can sign up to explore its capabilities further.
Dec 01, 2025
612 words in the original blog post.