Home / Companies / Datadog / Blog / July 2025

July 2025 Summaries

27 posts from Datadog

Filter
Month: Year:
Post Summaries Back to Blog
In the July episode of "This Month in Datadog," the focus shifts to the individuals behind key Datadog products, with Jeremy hosting discussions featuring Tristan Ratchford and Kevin Hu. Ratchford delves into Bits AI SRE, an advanced incident response model designed to alleviate the burden on on-call engineers by investigating alerts, coordinating incidents, and learning from each response. Meanwhile, Hu highlights Datadog Data Observability, which provides comprehensive visibility into datasets throughout their lifecycle, addressing issues before they escalate. Hu also shares insights about his journey from leading Metaplane to building innovative solutions at Datadog. Additionally, the episode introduces an AI voice interface aimed at accelerating incident response. As always, viewers are encouraged to subscribe to the YouTube channel for updates and explore the blog posts and release notes for a deeper understanding of the latest features and developments.
Jul 31, 2025 351 words in the original blog post.
In Kubernetes environments, applications commonly communicate with the Datadog Agent to transmit telemetry data using DogStatsD and Datadog APM, with communication modes managed by the Datadog Cluster Agent's Admission Controller. Traditionally, utilizing Unix domain sockets (UDS) for this purpose is preferred due to better performance and speed, but it requires mounting socket files via hostPath volumes, which conflicts with non-privileged Pod Security Standards (PSS). To address this, Datadog has introduced a Container Storage Interface (CSI) driver that allows UDS sockets to be mounted into pods using CSI volumes instead, ensuring compatibility with all PSS levels while maintaining security. The CSI driver supports multiple mount types for observability sockets, integrates with the Datadog Admission Controller for seamless adoption, and future updates plan to enhance support for APM Single Step Instrumentation libraries. This CSI-based method allows Kubernetes users to leverage UDS-based observability efficiently and securely within environments that adhere to strict security standards.
Jul 30, 2025 907 words in the original blog post.
In modern cloud environments, observability is crucial for SRE and DevOps teams to maintain insight into the real-time state of their systems, especially during outages or disruptions. Datadog has addressed this challenge by introducing Datadog Disaster Recovery (DDR), an active-passive observability solution that allows organizations to maintain visibility and telemetry data continuity by failing over to an alternate Datadog site when their primary site is affected. This approach offers a cost-effective alternative to the active-active model, which, while highly resilient, is expensive and operationally complex. DDR enables users to pre-configure their systems to redirect telemetry data to a secondary site, ensuring observability without the need for a costly and complex active-active setup. During a failure, customers can trigger a failover to a pre-established secondary site using Datadog's automation capabilities, ensuring continuous observation of their systems. This solution not only helps in maintaining observability continuity during outages but also aligns with regulatory compliance goals, allowing organizations to manage distributed services effectively without losing visibility.
Jul 30, 2025 1,194 words in the original blog post.
The text outlines Datadog's efforts to achieve FedRAMP High authorization, which necessitates complying with enhanced security controls beyond the FedRAMP Moderate baseline. To meet these requirements, Datadog has utilized its out-of-the-box capabilities for monitoring, logging, and compliance within its platform, leveraging tools such as Cloud SIEM, Incident Management, and Cloud Security Misconfigurations. These tools help address specific control families like Audit and Accountability, System and Information Integrity, and System and Communications Protection, ensuring robust security measures for federal data. The process involves using custom dashboards to track compliance initiatives, including implementing FIPS-validated cryptography and automating user account reviews to meet stricter FedRAMP High standards. The text also highlights the importance of continuous visibility into compliance efforts and the use of Datadog's platform for monitoring highly sensitive data in alignment with government standards.
Jul 30, 2025 2,219 words in the original blog post.
The text discusses the increasing adoption of AI and the associated emergence of tools like the Model Context Protocol (MCP), which facilitates the integration of applications with large language models (LLMs). MCP servers act as intermediaries between hosts and a range of data sources and services, making them susceptible to various security threats, including prompt injections, tool poisoning, and misconfigurations. The text highlights the potential vulnerabilities in MCP server interactions, such as unauthorized actions due to malicious tool definitions and the risks posed by third-party servers with inadequate security measures. It emphasizes the importance of monitoring MCP server activity and configurations, as well as LLM input and output, to detect and mitigate security risks. Additionally, the text underscores the need for robust authentication and authorization practices and suggests using tools like Datadog to enhance visibility and security of MCP deployments.
Jul 28, 2025 1,646 words in the original blog post.
Continuous profiling is transforming how developers monitor and optimize software performance by enabling non-stop, low-overhead data collection on application behavior in production environments. Unlike traditional profiling methods that were time-consuming and resource-intensive, continuous profiling integrates seamlessly with other telemetry data, such as metrics, logs, and traces, offering a comprehensive view of application performance. This approach helps teams quickly diagnose and resolve issues, improve release velocity, and enhance the customer experience. Its adoption is gaining momentum, with OpenTelemetry recognizing it as a critical observability signal alongside metrics, logs, and tracing. Companies like Datadog have leveraged continuous profiling to achieve significant cost savings and operational efficiencies, underscoring its value as an essential tool in today's competitive software landscape.
Jul 25, 2025 1,537 words in the original blog post.
The text outlines the complexities organizations face when transitioning from a permissive to a restrictive egress traffic policy in Kubernetes environments. Initially, organizations often allow open internet access for agility, but as their Kubernetes use grows, the need for stringent security measures becomes paramount. Tools like Cilium and AWS VPC CNI can help enforce deny-by-default policies, but implementing these without disrupting services is challenging. Datadog's Cloud Network Monitoring (CNM) assists by providing detailed insights into network traffic, enabling organizations to identify which Kubernetes namespaces require internet access. The text explains how to use CNM to gather traffic data, create lists of namespaces based on their internet access needs, and apply targeted network policies to restrict egress traffic effectively. This ensures a secure environment while maintaining necessary connectivity, using either Cilium or AWS VPC CNI for policy implementation.
Jul 23, 2025 2,510 words in the original blog post.
The Datadog Summit in San Francisco on September 9, 2025, serves as a vibrant celebration of the Datadog community, bringing together engineers, SREs, developers, and partners to share knowledge and innovations. The event kicks off with a keynote from Datadog executives and product managers, highlighting the latest product updates and real-world applications of their technology. The afternoon features a fireside chat and breakout sessions where Datadog engineers and customers discuss technical solutions and lessons learned. Attendees can participate in hands-on workshops led by Datadog experts, focusing on Kubernetes monitoring, OpenTelemetry, software quality with APM and distributed tracing, and building a small LLM application. The summit also provides networking opportunities through product demos, gamified experiences, and a reception with refreshments. A day prior, the Datadog Partner Network hosts an exclusive Partner Day to discuss strategic growth and collaboration with current and prospective partners through guided roundtable discussions.
Jul 22, 2025 398 words in the original blog post.
Candace Shamieh's piece delves into the roles of Bella Barbera and Emilio Rodriguez, two Customer Success Sales Engineers (CSSEs) at Datadog who transitioned from nontraditional backgrounds into sales engineering, highlighting their contributions to the company's Sales Engineering team. Bella, previously a mechanical engineer, shifted her career focus to software by enhancing her technical skills through self-directed learning, ultimately finding her niche in the collaborative and problem-solving aspects of sales engineering. Emilio, initially a Customer Success Associate, sought to deepen his technical knowledge to provide more comprehensive customer support, leveraging mentorship and self-directed study to transition into his current role. Both individuals exemplify Datadog's culture of curiosity, continuous learning, and cross-functional collaboration, contributing to customer success and the growth of the Sales Engineering team through leadership, innovative problem-solving, and cross-cultural engagement. Their journeys underscore the value Datadog places on initiative and adaptability, encouraging others to consider careers in a dynamic, supportive environment.
Jul 17, 2025 1,234 words in the original blog post.
In early 2025, the release of Go 1.24 introduced the Swiss Tables map implementation, promising reduced CPU and memory usage, but unexpectedly resulted in a 20% increase in memory consumption across several environments. Upon investigation, the increased memory usage was traced to a subtle regression in the memory allocator introduced by a runtime refactor, which caused more virtual memory to be committed to physical RAM, leading to higher resident set size (RSS) without affecting Go's internal metrics. The root cause was identified as the removal of an optimization in the mallocgc function, which previously avoided unnecessary zeroing of memory when allocating large objects containing pointers. This oversight was rectified with the help of the Go community, and a fix was implemented to restore the optimization in Go 1.25. The investigation not only resolved the issue but also revealed that in high-traffic environments, the new Swiss Tables implementation significantly reduced memory usage, demonstrating the complex interplay between runtime changes and memory management.
Jul 17, 2025 1,848 words in the original blog post.
The text discusses the challenges organizations face when migrating legacy web applications to serverless environments like AWS Lambda, which are event-driven and stateless, unlike traditional persistent HTTP servers. The AWS Lambda Web Adapter facilitates this transition by enabling web server-based applications to run on Lambda with minimal changes, allowing organizations to leverage Lambda's scalability and cost advantages. However, this setup can obscure the request life cycle and complicate the connection of logs, traces, and metrics without additional instrumentation. Datadog's integration with the Lambda Web Adapter addresses these challenges by providing full visibility into real-time metrics, traces, and logs from Lambda-hosted web applications. The integration involves including Datadog's Lambda Extension and Web Adapter binary in the Lambda function package and configuring necessary environment variables. This setup allows Datadog to automatically capture telemetry data, enhancing monitoring, performance detection, and troubleshooting capabilities without needing to modify application code. Datadog extends its Serverless Monitoring to include comprehensive insights into web applications running through the Lambda Web Adapter, capturing enhanced metrics like cold start durations and memory usage, and correlating logs with traces and metrics for unified visibility into runtime activity and application errors.
Jul 17, 2025 617 words in the original blog post.
The upgrade to Go 1.24 initially introduced a memory regression that increased physical memory usage across Datadog's services, prompting a collaboration with the Go community to identify and fix the issue. Surprisingly, the new update led to a significant decrease in memory usage in high-traffic environments, driven by the new Swiss Tables implementation that improved memory efficiency for large in-memory maps like the shardRoutingCache. By examining live heap profiles, it was discovered that Swiss Tables allowed for faster probing and eliminated the need for overflow buckets, yielding substantial memory savings. Further reductions in memory usage were achieved by optimizing the data structure, particularly by removing unused fields and using smaller data types. These changes not only mitigated the initial regression but also enhanced overall memory efficiency, demonstrating the importance of detailed runtime metrics and the potential for small optimizations to yield significant performance improvements.
Jul 17, 2025 3,813 words in the original blog post.
Datadog is hosting a summit in Sydney on August 19, 2025, celebrating its community of engineers, developers, and SREs who contribute to the enhancement of the platform. The event will feature a keynote by Mitch Ward, Datadog's Director of Engineering, who will discuss the new Australian data center and recent product updates. Attendees can participate in three hands-on workshops focusing on securing cloud-native architecture, building a tiny LLM application, and monitoring user flows. The summit also includes networking opportunities, product demos, and a closing reception to facilitate community engagement and knowledge sharing among Datadog users.
Jul 16, 2025 247 words in the original blog post.
The text discusses the significance of container image metadata, such as digest, size, and created_by, which are crucial for debugging, optimizing, and managing security risks in containerized environments. Missing metadata fields can disrupt these processes, as they hinder the ability to trace image lineage, identify optimization opportunities, and correlate vulnerabilities with specific commands. The reasons for missing metadata include limitations of certain build tools like Jib and Crane, the nature of temporary intermediate layers, and issues with base images. The text provides solutions, such as upgrading tools and explicitly setting metadata fields, to address these challenges. It also emphasizes the importance of understanding and troubleshooting missing metadata to maintain effective observability and security, highlighting tools like Datadog Container Monitoring for enhanced visibility into container images.
Jul 16, 2025 1,429 words in the original blog post.
The OpenTelemetry (OTel) Collector serves a crucial function in managing telemetry data, with various distributions available to suit different production needs. While the otelcol-contrib distribution is popular for demos due to its comprehensive component library, it is not recommended for production environments because it includes unnecessary components that can increase deployment complexity and risk. Instead, organizations are advised to choose leaner, more specialized distributions, such as the core OpenTelemetry Collector, or opt for vendor-maintained distributions that offer better integration with specific backends but may lack customization. Alternatively, building a custom Collector using the OpenTelemetry Collector Builder (OCB) allows for precise component selection but requires significant maintenance. Despite the convenience of otelcol-contrib, the community recognizes the need for more straightforward, customizable build processes, highlighting ongoing discussions about improving OCB to create a better default option for users. Ultimately, selecting the right distribution depends on individual architectural needs and use cases.
Jul 16, 2025 1,988 words in the original blog post.
The integration of Datadog LLM Observability with Amazon Bedrock Agents provides enhanced monitoring and management capabilities for agentic AI applications that utilize large language models (LLMs) to perform complex, multi-step tasks. By capturing detailed telemetry data, this integration allows developers to track and optimize agent workflows, ensuring performance, reliability, and cost-effectiveness while addressing issues like latency, errors, and token usage. It also facilitates the evaluation of output quality and tool selection to verify task completion accuracy and safety, detecting issues such as prompt injections and toxic content. This comprehensive observability enables developers to troubleshoot and resolve issues swiftly, maintaining high standards for responsible AI use. With over a decade of experience integrating with AWS services, Datadog offers a robust solution for teams looking to build and monitor their agentic AI applications effectively.
Jul 15, 2025 742 words in the original blog post.
As applications grow in complexity, identifying the root causes of issues becomes challenging, particularly when monitoring strategies separate frontend and backend data. To effectively troubleshoot, it's crucial to have visibility into how different parts of the app interact. Synthetic monitoring and distributed tracing are tools that offer different perspectives; the former simulates user interactions to identify frontend bottlenecks, while the latter tracks backend requests. Integrating these tools can provide a comprehensive view of the system, enabling proactive issue resolution before production. Datadog facilitates this integration by merging tracing with synthetic monitoring, allowing teams to create code-free tests that assess app availability, performance, and usability across various platforms. This integration streamlines the troubleshooting process by displaying traces alongside errors and warnings, enhancing the ability to quickly identify and resolve issues. By combining these tools, developers can efficiently transition from identifying user experience issues to investigating their root causes, ultimately improving the development process.
Jul 15, 2025 1,072 words in the original blog post.
DASH 2025, Datadog's largest event to date, took place at the North Javits Center in New York City, bringing together thousands of attendees for a comprehensive program of keynotes, sessions, and workshops. The event featured significant product announcements such as Flex Logs, Bits AI, and Datadog IDP, aimed at enhancing observability and developer experiences. Prominent speakers, including Bhawna Singh and Dave Asai, shared insights into innovative uses of Datadog's technology, while actor Kumail Nanjiani provided a humorous keynote on resilience and change. The event also included over 100 sessions, with notable contributions from companies like Coinbase and Block, and hands-on workshops covering topics like Kubernetes and DevSecOps. Additionally, the Women in Tech Lunch Panel offered career advice from senior leaders, and the Partner Summit highlighted the crucial role of partnerships in Datadog's growth, culminating in the Partner of the Year Awards. The event underscored Datadog's commitment to expanding its global user community with regional programs and promised a return to New York City for DASH 2026.
Jul 15, 2025 795 words in the original blog post.
Modern websites increasingly depend on third-party applications and open-source tools to enhance functionality, but this reliance poses security and privacy risks, exposing them to sophisticated attacks like Magecart and web skimming. To address these vulnerabilities, Datadog has partnered with Reflectiz, a web exposure monitoring platform, to help organizations monitor their client-side stack and detect threats in real-time. By integrating with Datadog, Reflectiz provides security and privacy alerts, allowing users to investigate and mitigate risks directly from the Datadog platform. Organizations can track and benchmark their web exposure risk using Reflectiz's Exposure Rating system, which evaluates risks based on page types and other factors. This integration offers increased visibility, enabling security teams to focus on reducing exposure and improving risk posture by learning from both internal data and industry trends. Users can begin using the Reflectiz integration by purchasing a license or starting a free trial in the Datadog Marketplace.
Jul 11, 2025 871 words in the original blog post.
Datadog has been recognized as a Leader in the 2025 Gartner® Magic Quadrant™ for Observability Platforms for the fifth year in a row, highlighting its comprehensive approach to observability, security, and AI integration. The platform supports organizations in managing complex systems through features like end-to-end application performance management (APM), Bits AI for autonomous incident response, and LLM Observability for monitoring AI applications. Its unified platform encourages collaboration across IT operations, development, security, and business teams by providing a single source of truth, reducing mean time to resolution (MTTR), and enhancing root cause analysis. Datadog's commitment to customer feedback and ongoing partnerships with AI-native technologies underscores its dedication to helping users effectively observe, secure, and act across their technology environments. The recognition reflects Datadog's ability to support modern application development and AI adoption at scale, with OpenTelemetry support for standardized telemetry data collection.
Jul 10, 2025 583 words in the original blog post.
The text discusses the concept of data lineage, which is the metadata that traces the flow and transformation of data within data pipelines, crucial for ensuring data quality, security, and compliance. It highlights the importance of data lineage in understanding and managing complex data ecosystems, facilitating data provenance, and ensuring regulatory compliance. The text introduces OpenLineage, an open-source framework that provides a standardized approach to lineage collection across different platforms, making it easier to track data movement and transformations. The OpenLineage core model uses a JSON Schema to define key types of lineage metadata, such as datasets, jobs, and runs, allowing for extensibility through customizable facets. The text also illustrates how lineage data is collected, stored, and visualized in a typical data platform, emphasizing the role of OpenLineage in enhancing data observability and interoperability in data pipelines.
Jul 09, 2025 1,505 words in the original blog post.
The Datadog mobile app is designed to optimize incident response by providing on-call engineers with real-time access to critical information, helping them manage alerts and mitigate issues from anywhere, at any time. By delivering instantly actionable push notifications, the app allows users to receive high-priority alerts that can override phone settings like "Do Not Disturb," while lower-urgency alerts are delivered quietly to reduce fatigue. The app's Notification Center consolidates alerts, enabling users to track and manage them efficiently, while integrated features like Bits AI SRE assist in identifying root causes and suggesting remediation steps. The app enhances coordination by offering visibility into team schedules, rotations, and escalation policies, ensuring that responders remain informed and aligned. Widgets for active incidents and other key metrics provide at-a-glance information, fostering situational awareness without needing to open the app. As a result, the Datadog mobile app helps teams reduce their mean time to repair by streamlining communication and response efforts, making it a valuable tool for maintaining operational agility.
Jul 09, 2025 849 words in the original blog post.
Datadog Error Tracking, integrated with GitHub, enhances the debugging process for developers by providing detailed insights into errors, facilitating quicker resolution. By automatically grouping similar errors and offering contextual metadata, it aids in identifying the exact lines of code responsible for an error and highlights the commit likely to have introduced it. The integration utilizes GitHub CODEOWNERS to clarify code ownership, ensuring that the correct teams are accountable and can prioritize issues effectively. Additionally, the feature of Suspect Commits allows for automatic assignment of issues to the developers responsible for the problematic changes, minimizing manual intervention and accelerating the triage process. This comprehensive approach reduces investigation time, keeps developers focused, and improves the overall efficiency of error resolution in complex codebases.
Jul 09, 2025 955 words in the original blog post.
Organizations often face challenges when onboarding security teams to a Security Information and Event Management (SIEM) system, primarily due to issues with data ingestion, storage costs, and integration capabilities. Datadog Cloud SIEM addresses these challenges by providing flexible data ingestion, over 900 prebuilt integrations, and customizable APIs that allow organizations to maintain their tools and workflows without disruption. Its Observability Pipelines enable efficient log management by filtering, deduplicating, and transforming logs, reducing costs, and enhancing compliance with privacy standards. The platform also offers Content Packs to streamline onboarding, providing bundled threat detection rules, dashboards, and automated workflows for rapid adoption. By simplifying log management and enhancing integration, Datadog empowers organizations to modernize their security operations efficiently and effectively.
Jul 03, 2025 1,275 words in the original blog post.
The special episode of "This Month in Datadog" highlights the events and announcements from DASH, Datadog's annual conference, where a variety of new products and features were revealed to enhance AI monitoring, system reliability, and developer experience. Key innovations include Datadog GPU Monitoring, AI Agents Console, Bits AI SRE, AI voice interface, and APM Latency Investigator, among others. Additionally, the episode features the Datadog Ambassadors program, aimed at creating a global network of technical experts who engage in solving problems, sharing knowledge, and contributing to open source projects. The program's participants were actively involved in community sessions and talks during DASH 2025, focusing on topics like breaking down monolithic architectures. Viewers are encouraged to explore Datadog's new offerings through the company's platform or a free trial and to stay updated with upcoming episodes and blog posts.
Jul 03, 2025 336 words in the original blog post.
The text outlines how Datadog's tools, including App Builder, Workflow Automation, and Datastore, streamline the process of scaling systems by consolidating multiple platforms into a single, user-friendly interface. These tools facilitate efficient management of cloud resources and improve communication within organizations by automating routine tasks and integrating third-party services. In particular, Datadog's product management team has used these tools to enhance their user outreach process, reducing context switching and manual data handling. By leveraging Datadog's datastore, workflows, and app functionality, they have created a consolidated app that combines user analytics, communication features, and AI integration, ultimately enabling more effective and personalized user interactions. This approach not only saves time but also shifts focus from repetitive tasks to more strategic activities.
Jul 03, 2025 1,110 words in the original blog post.
Datadog is advancing its capabilities for public-sector IT teams by achieving "In Process" status for GovRAMP High Authorization, which builds upon its existing FedRAMP Moderate and High authorizations. This development allows state, local, and education (SLED) institutions to utilize Datadog's unified observability platform to meet stringent security standards and improve digital service delivery. GovRAMP serves as a shared security framework for evaluating cloud service providers, focusing on continuous monitoring, vulnerability remediation, and independent audits based on NIST 800-53 Rev. 5 controls. By obtaining GovRAMP High status, Datadog enables SLED teams to enhance government service delivery with real-time visibility into complex IT environments, improving system performance and reducing operational costs. A case study involving a state transportation department illustrates how Datadog's platform can prevent traffic management disruptions by providing real-time insights and AI-powered anomaly detection. Current and potential Datadog customers in the public sector can benefit from enhanced security and compliance capabilities by utilizing this platform.
Jul 01, 2025 612 words in the original blog post.