Home / Companies / Datadog / Blog / April 2026

April 2026 Summaries

33 posts from Datadog

Filter
Month: Year:
Post Summaries Back to Blog
The Anomaly Reasoning Framework Benchmark (ARFBench) is a newly introduced benchmark designed for time series question-answering (TSQA) tasks, derived from real internal incidents at Datadog using its telemetry data. ARFBench aims to evaluate the performance of AI models, such as large language models (LLMs), vision-language models (VLMs), and time series foundation models (TSFMs), in diagnosing system anomalies by analyzing observability metrics. It highlights the substantial room for improvement in current models and introduces a novel hybrid TSFM-VLM model, Toto-1.0-QA-Experimental, which demonstrates promising results by achieving high accuracy and F1 scores while offering efficiency gains. The benchmark is structured into three tiers of increasing difficulty, emphasizing compositional reasoning and the integration of context across data modalities. ARFBench sets a new superhuman frontier when combined with human expertise, showcasing complementary strengths between models and experts. The framework is positioned as a significant step in developing end-to-end agentic systems for incident response, with resources available on platforms like Hugging Face and GitHub for further exploration and development.
Apr 24, 2026 2,015 words in the original blog post.
In exploring effective network tests for applications, the text emphasizes the importance of selecting appropriate protocols such as TCP, UDP, and ICMP to reflect actual user experiences and identify network issues. Network tests, particularly those using traceroute queries, can help determine whether problems originate from the network or application layer by measuring latency and availability across network paths. The text highlights the functionality differences between TCP and UDP, with TCP offering reliable data delivery through features like retransmission and connection state tracking, while UDP provides faster transmission suitable for high-speed applications like video streaming. Datadog Synthetic Monitoring is presented as a tool that facilitates the creation and scheduling of network tests, enabling users to simulate realistic traffic and visualize results for comprehensive network analysis. By using these tests, users can address network issues efficiently without misattributing them to application problems, ensuring network stability and performance reliability.
Apr 23, 2026 1,373 words in the original blog post.
Product managers are often perceived as the CEOs of their products, but the reality is that they face significant challenges in obtaining timely and comprehensive insights from various product signals. These signals—categorized as pulse (immediate performance data), pain (user frustration indicators), and proof (long-term business metrics)—arrive at different latencies, which can hinder swift decision-making and product iteration. While pulse and pain signals can provide immediate insights into product issues, proof signals typically take longer to gather, often resulting in lost opportunities for improvement by the time they are analyzed. The key to enhancing product growth lies in understanding and leveraging these signal latencies, enabling teams to act on faster signals like pulse and pain while waiting for longer-term proof metrics. The integration of tools such as session replays and warehouse-native metrics can provide a more cohesive view of product impact, allowing for efficient experimentation and quick iterations. By consolidating these signals, product managers can shift from coordinating information to directly piloting product decisions, ultimately leading to more confident and timely product development.
Apr 23, 2026 1,081 words in the original blog post.
Datadog is enhancing its collaboration with Google Cloud to address challenges in managing the full AI stack, from data pipelines and infrastructure to security operations, by providing a unified platform for AI application teams. This platform offers tools for evaluating and troubleshooting AI applications, optimizing costs and performance on GPUs and TPUs, improving data reliability, and strengthening security with AI-driven investigation and response capabilities. Datadog's solutions, including LLM Observability and GPU Monitoring, aim to streamline visibility into agent behavior, data health, and infrastructure performance, while Bits AI Security Analyst assists in threat detection and response. By integrating with Google Cloud's infrastructure, Datadog helps teams reduce complexity, improve reliability, and move faster, positioning itself as a key partner for AI success on Google Cloud.
Apr 22, 2026 867 words in the original blog post.
Cloud environments produce an overwhelming number of security signals daily, leading security engineers and analysts to spend excessive time triaging rather than addressing genuine threats. Automating the manual aspects of cloud security investigations, such as identifying related signals and building timelines, allows teams to focus on tasks requiring human judgment. The article explores automating such processes with Bits AI Security Analyst, which streamlines behavioral analysis and threat correlation. Using a specific scenario involving an AWS IAM user suddenly generating numerous API calls, it illustrates how Bits AI identifies suspicious activity and assesses potential threats by analyzing signal patterns, API call structures, and threat intelligence data. The scenario highlights the importance of distinguishing between legitimate and suspicious actions, such as phishing campaigns, by examining API call patterns, user agents, and IP addresses. The investigation process involves identifying potential account compromises through unauthorized API calls and determining the actor's intentions, such as phishing, by examining the structure and frequency of these calls. Ultimately, Bits AI Security Analyst aids in accelerating cloud security investigations, enabling analysts to make informed decisions and prioritize containment and remediation efforts based on concrete behavioral and threat intelligence data.
Apr 22, 2026 1,936 words in the original blog post.
Organizations often struggle to efficiently gather and utilize developer feedback due to its dispersion across various disconnected systems, which complicates the correlation of feedback with performance outcomes. Datadog addresses these challenges by integrating Datadog Forms and Datadog Sheets into the existing operational environment, allowing for structured feedback collection and analysis without leaving the platform. Forms enable the creation and distribution of surveys directly within Datadog, supporting various data types and ensuring consistent responses, while Sheets facilitate the analysis of feedback alongside other data sources, enriching it with metadata and visualizing trends over time. This integration helps organizations move from fragmented input to clear, actionable insights, allowing them to identify patterns and make informed decisions to improve internal platforms, tooling, and processes.
Apr 22, 2026 1,099 words in the original blog post.
UK organizations are increasingly required to consider data residency requirements, ensuring operational data remains within national boundaries, particularly in regulated sectors like financial services, healthcare, and the public sector. To address this, Datadog plans to launch a UK availability zone in partnership with AWS by late 2026, enabling organizations to store and process telemetry data within the UK. This initiative ensures that observability data generated by AWS workloads is stored in-region without altering workflows or reducing access to Datadog's full platform, thereby supporting compliance with UK data protection regulations and maintaining visibility for complex, distributed systems. The UK availability zone aims to reduce latency, simplify audit workflows, and provide end-to-end visibility, catering to the needs of organizations managing cloud-native architectures and AI-intensive workloads, without compromising operational resilience or data governance. This aligns with similar efforts in other regions, such as Australia's Sydney availability zone, and reflects a broader trend of integrating local data hosting into system architecture from the outset.
Apr 21, 2026 965 words in the original blog post.
Developers and Site Reliability Engineers (SREs) using Microsoft Azure DevOps often encounter fragmented workflows and delays in issue detection, particularly in troubleshooting errors and assessing code quality. Datadog's Azure DevOps Source Code integration addresses these challenges by integrating repositories directly into Datadog, providing a unified, code-aware view that enhances Application Performance Monitoring (APM), CI visibility, and code security. This integration allows teams to analyze code health and identify vulnerabilities earlier in the software delivery life cycle (SDLC) without redesigning pipelines. It links trace data to underlying code, offering contextual insights that streamline the path from error detection to resolution. Automated feedback in pull requests highlights issues such as vulnerabilities and test regressions, enabling quick assessments and direct remediation within Azure DevOps. Additionally, PR Gates enforce consistent quality and security standards before code merges, ensuring a higher baseline of code quality across repositories. By centralizing these processes, the integration simplifies workflows, reduces tool switching, and accelerates issue resolution, as detailed in the Azure DevOps Source Code integration documentation.
Apr 21, 2026 941 words in the original blog post.
Datadog has developed an innovative system to embed widget metadata invisibly within screenshots, using a pixel-level encoding scheme that remains resilient across various color profiles and display densities. This approach allows users to share visualizations from Datadog dashboards as interactive widgets, even when pasted into platforms like Slack. Traditional screenshots, while quick and easy, often miss critical context such as time range and underlying queries. The new system uses watermarking techniques, including fuzzy encoding, to embed metadata into widget borders without affecting user experience or performance, allowing over a billion watermarks to be created daily. This capability not only preserves the original data attributes but also enables the reconstruction of widgets from screenshots, facilitating a seamless transition between static images and live data. The system also incorporates backend processes to cache and retrieve metadata efficiently, while frontend optimizations ensure minimal impact on app responsiveness. By enhancing the way visual data is shared and used, Datadog is pushing the boundaries of data visualization and sharing within collaborative environments.
Apr 21, 2026 4,405 words in the original blog post.
As organizations expand, managing observability efforts becomes increasingly complex, with more teams leading to a proliferation of dashboards, monitors, API keys, and custom configurations. This often results in governance becoming a reactive process focused on mitigating waste rather than proactively establishing standards. Datadog's Governance Console addresses this challenge by offering a centralized interface that transforms configuration and usage data into actionable insights, automating best practice enforcement. It allows organizations to monitor Datadog product adoption and configuration, preventing configuration drift and enhancing accountability. By providing a consolidated view of organizational usage, it helps answer strategic questions, identify gaps in governance standards, and prioritize cleanup efforts. The console's per-product insights and catalog of controls enable administrators to optimize platform feature adoption, reduce waste, and enforce organizational standards, ultimately managing observability at scale with reduced manual effort.
Apr 17, 2026 817 words in the original blog post.
Experiments, particularly randomized controlled trials, have proven to be timeless tools for organizations seeking to understand causation and improve decision-making, regardless of technological advancements. These trials, commonly associated with A/B testing, are utilized by leading companies like Meta, Google, and Amazon to run tens of thousands of experiments annually, thereby refining product features, user interfaces, and deployment strategies. Notably, canary testing has become crucial in software deployment, enabling organizations to test changes on a small scale before full implementation, preventing incidents like the significant outage faced by Google Cloud in 2025. Companies such as Microsoft, Amazon, and Duolingo leverage experimentation to enhance metrics, from reducing latency to improving user engagement, while AI models are tuned and evaluated through online evaluations. This iterative process drives a learning velocity that builds a competitive advantage, as organizations evolve technically and culturally to respond swiftly to market changes. Ultimately, experimentation is not just about success or failure but about learning and adapting quickly, offering insights that compound over time to create substantial value.
Apr 17, 2026 1,859 words in the original blog post.
Organizations investing in AI are facing challenges with rapidly increasing telemetry data volumes and associated costs, leading them to adopt more efficient data pipelines like Observability Pipelines. These pipelines enable teams to collect, enrich, transform, and route telemetry data to various destinations, such as ClickHouse and preferred Security Information and Event Management (SIEM) systems, offering flexibility and high performance. While tools like the OpenTelemetry (OTel) Collector standardize data collection, they lack routing flexibility and enrichment capabilities. Observability Pipelines address this gap by allowing data to be collected from both OTel and non-OTel sources and routed to different environments, including hybrid and multi-cloud setups. This approach helps standardize data formats, which is crucial for consolidating telemetry into fast query layers like ClickHouse, ensuring data quality and easy analysis. The system also supports compliance efforts through features like sensitive data redaction and offers monitoring tools to track pipeline performance and health. By integrating Datadog's Observability Pipelines, teams can manage data flows effectively, maintaining insights across diverse systems and preventing issues like data loss and processing bottlenecks.
Apr 16, 2026 1,454 words in the original blog post.
Sarjeel Yusuf's Single Step Instrumentation (SSI) enhances Datadog's Application Performance Monitoring (APM) by automatically discovering and instrumenting services across hosts, providing teams with comprehensive visibility with minimal setup. As environments expand, teams often seek more control, which is where instrumentation rules offer a solution by allowing users to specify which services generate traces, thus optimizing debugging and performance decision-making. These rules, applicable on both Linux and Windows, enable users to manage instrumentation centrally and exclude non-essential services, such as batch jobs that do not contribute to performance analysis, through attributes like operating system, process details, and runtime characteristics. By leveraging SSI and instrumentation rules, teams can maintain efficient tracing while minimizing unnecessary data and costs, with the flexibility of centralized or local configuration deployment.
Apr 16, 2026 785 words in the original blog post.
In March 2026, a GitHub account named hackerbot-claw, self-described as an "autonomous security research agent," targeted open-source repositories by exploiting misconfigurations in GitHub Actions workflows, notably impacting a repository from Datadog. This incident highlighted the growing concern of AI agents autonomously discovering and exploiting CI/CD vulnerabilities. The campaign leveraged common misconfigurations such as unsafe pull_request_target configurations, unspecified workflow permissions, unpinned actions, and overprivileged tokens, which are prevalent in many public repositories. Datadog's Infrastructure as Code (IaC) Security tool can help mitigate these risks by scanning workflows for vulnerabilities before they are merged, ensuring issues are caught in the diff and blocking merges until resolved. The event prompted a comprehensive audit of Datadog's CI/CD security, leading to expanded coverage in areas like trigger and condition safety, as well as supply chain and runtime integrity. Implementing best practices, such as pinning actions to commit SHAs, setting explicit permissions, and avoiding the interpolation of user-controlled input into run blocks, can further secure GitHub Actions workflows.
Apr 16, 2026 1,311 words in the original blog post.
Datadog's App and API Protection (AAP) now provides in-process application security monitoring specifically for Python AWS Lambda functions, addressing the challenge of securing serverless environments where traditional perimeter defenses fall short. By integrating with the application runtime through Datadog tracing libraries, AAP offers deeper visibility into how requests interact with code, allowing detection of exploit attempts and correlation with vulnerable code paths. This approach enhances security by embedding AAP directly into the Python runtime, enabling real-time observation of function calls and request contexts. The Exploit Prevention feature adds runtime application self-protection (RASP) by analyzing data flow to detect injection attacks, while account takeover (ATO) attempts are identified through monitoring authentication-related events across invocations. This integration allows for a more comprehensive and responsive security strategy, helping teams quickly identify and mitigate threats in serverless applications.
Apr 14, 2026 791 words in the original blog post.
Offline evaluation is a crucial practice for developing reliable LLM-powered applications and agents, as it allows teams to test changes against known scenarios before deploying them to production. This method helps identify potential issues early on, reducing the risk of user frustration and revenue loss. Offline evaluation involves using curated datasets with annotated test cases that cover core use cases and potential edge cases, allowing developers to benchmark and iterate on AI agents efficiently. A robust evaluation framework includes data, tasks, and evaluators: data consists of annotated test cases, tasks involve the logic that produces outputs, and evaluators measure the quality of those outputs. Such a framework helps ensure that agents perform reliably by comparing different versions and catching regressions before they affect end users. Datadog's LLM Experiments provides tools for creating and managing these components, enabling developers to perform offline evaluations and improve AI agents with precision, scalability, and reduced incident rates.
Apr 14, 2026 2,807 words in the original blog post.
An AI-native static application security testing (SAST) tool has been developed to enhance vulnerability detection by utilizing large language models (LLMs) for more accurate, context-aware analysis compared to traditional SAST methods. This open-source solution, which scans code changes incrementally, aims to reduce false positives and improve detection rates for vulnerabilities such as SQL injection and cross-site scripting, as demonstrated by its superior performance on OWASP benchmarks. Although more costly due to multiple LLM calls, the tool mitigates expenses through incremental analysis, performing full repository scans only when necessary. The project is integrated within Datadog, allowing scalable operations for each code change and leveraging Datadog LLM Observability to monitor performance and costs. By open sourcing this tool, Datadog seeks to foster transparency and collaboration within the security community, aiming to further refine and expand its capabilities, including potential agentic scanning techniques for deeper contextual understanding. The AI-native SAST solution is available on GitHub, with incremental analysis currently previewed for Datadog customers.
Apr 10, 2026 1,201 words in the original blog post.
Engineers at Datadog encountered an unexpectedly high-latency index scan in a PostgreSQL table, despite using an index scan that is typically efficient. The issue was traced to a mismatch between the column order of the composite index and the query's filtering predicates, which caused unnecessary disk I/O and high node costs. By creating a targeted index that matched the query's predicate order, engineers reduced the query latency from over 300 ms to 38 μs, demonstrating the importance of aligning index structures with query patterns. Datadog's Database Monitoring (DBM) tool now includes enhancements to automatically detect and suggest fixes for suboptimal index scans across databases, helping users identify and resolve inefficiencies without manual investigation.
Apr 09, 2026 969 words in the original blog post.
Platform engineering teams often struggle to demonstrate measurable value, with over 40% of initiatives failing to do so within the first year, leading to risks such as defunding or deprecation. To accurately calculate a platform's return on investment (ROI), teams need to differentiate between metrics that measure platform effectiveness and those used for investigations. Metrics are categorized into a three-tier hierarchy: outcome, driver, and diagnostic metrics, each serving a distinct purpose. Outcome metrics provide a high-level view of platform health and are crucial for confirming platform value, while driver metrics identify bottlenecks, and diagnostic metrics investigate root causes. The DORA framework is highlighted as a leading standard for measuring software delivery performance, influencing organizational outcomes and team well-being. Datadog offers automated tools for tracking these metrics, providing actionable insights and enhancing the security posture of platforms. Common pitfalls include optimizing driver metrics in isolation, promoting diagnostic metrics to KPIs, and retaining stale metrics, all of which can mislead and hinder decision-making. By adopting a structured approach to metric selection and evaluation, platform teams can better demonstrate the organizational value of their investments and make informed, strategic decisions.
Apr 09, 2026 2,511 words in the original blog post.
Recorded Future's integration with Datadog enhances real-time threat intelligence by incorporating feeds of indicators of compromise (IOCs) such as malicious IP addresses, domains, and vulnerabilities, directly into Datadog's platform. This integration, which is the first of its kind for threat intelligence in Datadog, provides security teams with enriched context, including risk scores and threat associations, to prioritize responses more efficiently. By capturing Recorded Future's Classic Alerts and Playbook Alerts, the integration allows for seamless analysis alongside application and infrastructure data, making it easier to correlate external threat signals with internal activity. Datadog's Cloud SIEM uses this enriched intelligence to improve detection and prioritization of threats, thereby enabling faster, more informed responses without manual triage. The Recorded Future Content Pack further simplifies onboarding with prebuilt dashboards and out-of-the-box detection rules that help identify and prioritize threats, facilitating a strengthened security posture through the unified platform.
Apr 09, 2026 794 words in the original blog post.
Massimo Sporchia Boomi, an Integration Platform as a Service (iPaaS), is enhanced with native OpenTelemetry (OTel) support to improve observability of its integration runtimes, addressing previous challenges in gaining insights into runtime performance. By utilizing the industry-standard OpenTelemetry Protocol (OTLP), organizations can now export traces, logs, and metrics from Boomi Atoms without the need for third-party agents. To further enhance this observability, the platform can integrate with Datadog using the Datadog Distribution of the OpenTelemetry Collector (DDOT), offering comprehensive visibility into Boomi processes and their interactions with downstream services and database queries. This integration allows for the correlation of runtime telemetry with infrastructure metrics and JVM-level insights, turning Boomi from an operational blind spot into a transparent part of the application stack. The combination of Boomi’s OTel support and Datadog's capabilities enables operations teams to build real-time dashboards and gain end-to-end visibility into process execution, making it easier to diagnose and resolve issues effectively.
Apr 09, 2026 2,007 words in the original blog post.
Ingress NGINX, a key component for managing external traffic in Kubernetes, reached its end-of-life in March 2026, prompting an urgent need for organizations to migrate to the Kubernetes Gateway API due to security vulnerabilities like IngressNightmare and CVE-2026-24512. The Gateway API offers enhanced traffic management directly within its core resources, unlike Ingress NGINX, which relied on custom annotations. Migrating successfully requires a structured approach, starting with choosing a suitable Gateway API controller, capturing performance baselines, and installing the Gateway API controller alongside Ingress NGINX to ensure seamless transition. This involves translating existing Ingress configurations into Gateway and Route resources, verifying their acceptance, and gradually shifting production traffic while monitoring key metrics to detect any regressions. The migration not only involves technical changes but also organizational shifts, as it delineates clear roles and responsibilities between infrastructure providers, cluster operators, and application developers. The Gateway API's role-oriented design provides flexibility with advanced routing capabilities such as traffic splitting, hostname matching, and multi-protocol routing, making it a robust and modern solution for Kubernetes environments.
Apr 08, 2026 3,760 words in the original blog post.
In the continuation of a series on CI/CD security, this article applies threat modeling principles to GitHub as a source code management tool, examining historical attacks and discussing preventative measures and response workflows. By identifying inputs, identities, and risks associated with GitHub, the article highlights how unauthorized access, backdoor entry, data exfiltration, and malicious code execution can occur. It emphasizes the importance of detection methods using tools like Datadog Static Code Analysis, CodeQL, and Dependabot to identify and mitigate threats. The discussion includes real-world examples of attacks, such as the Shai-Hulud npm worms and unauthorized OAuth token access, illustrating how attackers exploit vulnerabilities in GitHub environments. The article also underscores the role of security tools like Datadog Cloud SIEM in detecting and preventing malicious activities. Additionally, it provides insights into safeguarding against compromised third-party dependencies by using static code analyzers and dependency checkers to detect vulnerabilities before they affect production environments.
Apr 08, 2026 1,843 words in the original blog post.
SCM and CI/CD pipelines are crucial in automating software delivery, yet they are vulnerable to attacks that can exploit these systems to gain unauthorized access, manipulate code, and deploy malicious software. The blog series leverages a threat matrix, adapted from the MITRE ATT&CK framework, to identify and map potential attack pathways specific to CI/CD systems, providing a structured approach to threat modeling. The series emphasizes the importance of integrating security practices and tools into CI/CD pipelines, focusing on securing the trust boundaries of SCM tools like GitHub, GitLab, and Bitbucket, as well as CI/CD tools such as Jenkins and GitHub Actions. It highlights how attackers can exploit vulnerabilities through compromised credentials and permissive access policies, leading to supply chain attacks. The blog offers guidance on detecting these threats using a CI/CD-specific threat matrix and outlines steps to secure environments, with a particular focus on GitHub, encouraging readers to employ proactive detection and response measures to safeguard their software delivery pipelines.
Apr 08, 2026 1,371 words in the original blog post.
AI-assisted development can expedite coding processes but also introduces heightened security risks, as it may inadvertently generate vulnerabilities, insecure dependencies, or expose secrets before human review. The Datadog Code Security MCP addresses these challenges by analyzing code in real-time, detecting and flagging issues such as SQL injection vulnerabilities, insecure dependencies, and hardcoded credentials as the code is written, which allows for immediate resolution before it reaches further stages like pull requests. This system consolidates various security checks into a single workflow by integrating static application security testing, software composition analysis, secrets detection, and infrastructure-as-code scanning, simplifying the developer's workflow and maintaining consistent security standards without the need for separate tools or reauthentication. The MCP server's local operation with a single authentication flow ensures ease of use, enabling developers to implement immediate security scans in their existing environments, and keeping security policies up-to-date with minimal effort. This solution is part of Datadog's comprehensive approach to securing AI-assisted development, offering tools like malicious pull request detection and AI Guard to protect software development lifecycles.
Apr 07, 2026 701 words in the original blog post.
Datadog's Bits AI SRE team developed a comprehensive evaluation platform to improve and trust their autonomous agent, Bits, which investigates production incidents by analyzing diverse data sources such as metrics, logs, and network telemetry. Initially, each feature added to Bits caused unforeseen regressions in other areas, highlighting the need for a robust evaluation system. The team built a replayable evaluation framework that uses curated labels representing real-world scenarios, allowing them to measure and refine Bits' performance effectively. This platform segments and scores investigations, tracks changes over time, and ensures that new features or models don't inadvertently degrade performance. The evolution of this system involved shifting from manual labeling processes to leveraging Bits itself for label creation and validation, thus increasing the label creation rate and quality. By embedding real-world noise into evaluations and automating much of the process, the team can now run extensive, realistic tests, preventing regressions and guiding development. This infrastructure not only supports Bits but also aids other Datadog teams in refining their agents, ensuring continuous improvement through systematic, large-scale evaluations.
Apr 07, 2026 3,378 words in the original blog post.
Tohn Furutani, an SRE Engineer at NTT DATA, discusses the evolution and challenges of implementing agentic AI systems, which are capable of making context-based decisions and executing multistep actions. He highlights the shift from single-use generative AI applications to more complex agentic systems that require not only accuracy and speed but also security, visibility, and consistent behavior in enterprise environments. To address these needs, NTT DATA adopted Amazon Bedrock AgentCore to standardize the execution platform and Datadog LLM Observability to enhance the visibility and evaluation of agentic actions and reasoning. This combination allows for better understanding and continuous improvement of the AI systems' decision-making processes. Through standardization and collaboration with AWS and Datadog, NTT DATA aims to create reusable solutions and share insights to support the adoption of operable and explainable agentic AI in enterprises.
Apr 07, 2026 1,602 words in the original blog post.
Datadog's Session Replay has introduced new capabilities for heatmap analysis, allowing teams to have more precise control over heatmap backgrounds by capturing specific UI states directly from their live sites. This innovation addresses the challenge of capturing dynamic or rare UI states, such as modals and post-login views, which were previously difficult to obtain through session replays. Users can now interact with their site in real-time to select the exact UI state needed, ensuring that heatmap data accurately reflects the user experience they aim to analyze. These custom backgrounds can be saved and shared across the organization, enabling team members to access a curated set of views without redundant efforts. The system allows for multiple backgrounds per heatmap view, enhancing flexibility and collaboration while maintaining a reliable and comprehensive library of user interactions.
Apr 06, 2026 534 words in the original blog post.
Bits Assistant is an innovative conversational interface integrated within Datadog's web and mobile applications, as well as collaboration tools like Slack, designed to streamline the process of finding and acting on data across various telemetry sources. It enables users to search, visualize, and take action using natural language, helping engineers quickly locate dashboards, monitors, logs, and traces without switching contexts. This tool assists in onboarding by providing in-context best practices and answers to platform questions, and it allows for the creation and modification of dashboards and notebooks through natural language prompts. By correlating signals and resolving issues across disparate data sources, Bits Assistant reduces the time needed to troubleshoot and document incidents, offering a seamless experience whether at a desk or on the go. This effectively bridges the gap between data and actionable insights, facilitating better collaboration and faster resolution of issues within the Datadog environment.
Apr 02, 2026 1,014 words in the original blog post.
Datadog's Session Replay feature provides teams with a video-like view of user interactions in their applications, which is enhanced by AI summaries and smart chapters to improve efficiency in identifying and analyzing user behavior. These AI summaries offer a concise overview of user sessions, highlighting key actions, outcomes, and areas of friction, thereby allowing engineers and product managers to quickly assess whether a replay is worth further investigation. Smart chapters break down sessions into distinct stages, making it easier to navigate and understand the user journey without manually scanning through long video recordings. This functionality supports various roles, such as engineers troubleshooting performance issues and product managers analyzing user experience problems, by integrating with Real User Monitoring (RUM) and Product Analytics to provide comprehensive insights. By streamlining the process of finding and interpreting relevant replays, these features enable teams to address technical failures and user experience issues more effectively, ultimately enhancing application performance and user satisfaction.
Apr 02, 2026 930 words in the original blog post.
Datadog has reimagined the on-call alert sound design in its mobile app, focusing on reducing alert fatigue and stress for engineers by employing empathetic and research-driven methods. Collaborating with audio specialists, the company developed a diverse library of notification sounds tailored to different user personas, such as those in noisy environments or shared spaces, individuals needing gentle wake-up calls, and those who appreciate humor in alerts. These sounds are designed to be effective without being harsh, prioritize the well-being of engineers, and support smoother cognitive transitions upon awakening. By emphasizing thoughtful sound design and strong alert management, Datadog aims to mitigate the human costs associated with on-call duties while maintaining the urgency necessary for effective problem resolution.
Apr 02, 2026 1,077 words in the original blog post.
Datadog Database Monitoring for ClickHouse, available in Preview, offers a comprehensive solution for understanding query performance and resource usage in both ClickHouse Cloud and self-hosted environments. It addresses the challenges engineers face in identifying memory-intensive and frequently running queries that can lead to performance issues. The tool provides a unified view of query activity, enabling teams to compare query patterns across clusters and prioritize optimization efforts based on overall system impact rather than just individual query latency. It aggregates metrics such as execution count, duration, memory usage, and I/O activity, helping to identify costly queries. Additionally, it captures and retains completed query samples, allowing for detailed post-incident analysis without the limitations of ClickHouse's own log retention settings. This database monitoring capability aims to help teams reduce costs, enhance performance, and effectively manage incidents by offering real-time and historical insights into query activity across all ClickHouse deployments.
Apr 02, 2026 852 words in the original blog post.
Datadog Experiments is a feature within the Datadog platform that enables product and engineering teams to conduct reliable experiments efficiently, directly integrating behavioral analytics, application performance telemetry, and warehouse-native business metrics. This tool addresses the challenges of traditional experimentation, which often involves fragmented workflows and delays due to dependency on data scientists and external analytics systems. By providing self-serve analysis tools and built-in guardrails, teams can define metrics, monitor real-time experiment results, and identify performance regressions early, facilitating a faster feedback loop and reducing the risk of costly restarts. Datadog Experiments also supports the seamless integration of warehouse-native business metrics, ensuring transparency and auditability while minimizing operational overhead. Additionally, the platform enables teams to optimize AI and LLM applications by combining offline evaluations with A/B testing, allowing for safe evaluation of changes in production. This unified approach helps teams make informed decisions based on comprehensive insights into the behavioral, performance, and business impact of product changes.
Apr 01, 2026 1,451 words in the original blog post.