Home / Companies / Datadog / Blog / September 2025

September 2025 Summaries

16 posts from Datadog

Filter
Month: Year:
Post Summaries Back to Blog
In the September episode of "This Month in Datadog," the focus is on new features and updates designed to enhance software development and network troubleshooting. The episode highlights tools like Datadog Feature Flags, which integrate flagging with observability for safer feature releases, and Datadog Network Path, which provides a clear visualization of network traffic to simplify the diagnosis of slowdowns. It also introduces Cloud Cost Management's integration with the Anthropic Usage and Cost Admin API, enabling teams to track Claude usage and costs efficiently. Additionally, the new Custom Processor in Observability Pipelines facilitates the migration of historical logs while maintaining cost control across various log formats. Other updates include enhanced debugging for failing pipelines, cloud cost recommendations for Azure and Google Cloud, and improved visibility into security risks with Cloudcraft and certificate chain issues via Windows Certificate Store integration. The episode encourages viewers to explore these features through blog posts and the Datadog platform, with an invitation to subscribe for future updates.
Sep 30, 2025 546 words in the original blog post.
Datadog Feature Flags streamlines the software release process by integrating feature flagging directly into the observability platform, allowing for seamless management of feature rollouts and health metrics in a single workflow. This integration eliminates the need for engineers to juggle multiple systems, reducing the risk of missing critical issues during feature deployment. With automated canary releases, rollbacks, and advanced experiments, teams can efficiently manage feature delivery, ensuring that any problems triggered by new features are swiftly addressed. The platform connects feature flags with telemetry data, enabling real-time monitoring of rollout impacts on metrics such as response times and error rates, which allows for quick identification and rectification of issues without additional deployments. By incorporating best practices like progressive rollouts and automated guardrails, Datadog Feature Flags supports safe experimentation and innovation, helping teams to release features quickly while maintaining reliability and reducing risk.
Sep 30, 2025 1,097 words in the original blog post.
Dan Green discusses how the Datadog Software Catalog enables software organizations to manage the complexity of their architectures by using custom entity types. These custom entities allow platform teams to define and manage components like pipelines, libraries, and AI agents, providing a clear visualization of dependencies and ownership across the ecosystem. This approach helps troubleshoot issues more efficiently by making hidden dependencies visible and assigning explicit ownership, ensuring accountability. Additionally, it allows for the application of specific best practices and governance across different component types, contributing to a more reliable and streamlined workflow. The catalog's unified view aids engineers in understanding component interactions, reducing the risk of missed connections, and enabling faster incident response by allowing immediate access to relevant documentation and contact information.
Sep 29, 2025 775 words in the original blog post.
In complex data pipelines, monitoring data lineage is crucial for ensuring data quality, regulatory compliance, and effectively troubleshooting issues. Data lineage provides a detailed map of data flows and transformations across various stages, enabling organizations to trace errors back to their upstream sources and assess their downstream impacts. Workflow orchestration tools such as Apache Airflow play a significant role in gathering lineage metadata, which helps in identifying root causes of errors, evaluating business impacts, and validating data quality with column-level precision. By utilizing lineage information, organizations can quickly address issues, maintain compliance, and ensure accurate data usage across different applications. Datadog Data Observability offers enhanced monitoring capabilities by integrating OpenLineage, allowing users to efficiently manage and resolve data pipeline errors, thereby improving their overall data governance strategy.
Sep 25, 2025 1,125 words in the original blog post.
The text outlines the integration of Site Reliability Engineering (SRE) and security teams to create a cohesive group focused on building platforms and products with a security-first mindset while maintaining established SRE practices. This integration addressed common operational pain points, such as tool duplication and gaps in audit logging, enhancing incident response through shared baselines, runbooks, and dashboards. These measures fostered a collaborative, blameless incident management culture, ensuring efficient problem-solving and risk mitigation. Regular cross-functional exercises, including security drills and chaos engineering, further refined incident management processes. The approach demonstrated that aligning team goals and improving process visibility can enhance collaborative incident response and platform resilience, offering a model for other organizations to consider.
Sep 19, 2025 1,742 words in the original blog post.
Datadog Synthetic Monitoring now supports Microsoft Active Directory with Kerberos single sign-on, enhancing its ability to proactively detect authentication failures that might disrupt employee access to critical applications and APIs. This integration allows for automated tests to be run from within a network's Active Directory domain, ensuring reliable authentication and minimizing operational risks. By continuously validating Kerberos-authenticated login workflows and API endpoints, Datadog Synthetics helps maintain productivity and reduces authentication-related outages. It also offers a unified monitoring approach across internal and external applications, improving the correlation of authentication failures with system metrics like latency and network performance. This new feature aims to close visibility gaps by providing a single source of truth for application performance, allowing early detection and swift resolution of potential authentication issues before they impact employees.
Sep 18, 2025 731 words in the original blog post.
Datadog has undertaken a substantial journey in cost-aware engineering by optimizing their infrastructure, saving $17 million, and developing Cloud Cost Management for customers. The company initially focused on manual code optimization of Go functions to reduce CPU usage in high-scale services. This manual process laid the groundwork for BitsEvolve, an internal agentic system for self-optimizing code that employs evolutionary algorithms to automate performance improvements. The system uses real-world observability data to identify optimization opportunities and runs continuous evaluation loops to refine code, ultimately leading to significant cost reductions. Despite the success of manual optimizations, scaling these efforts across the organization required transitioning to an automated system that integrates observability data, AI agents, and human expertise. This approach has not only improved specific functions by up to 90% but also aims to evolve into a self-optimizing system for broader performance-critical parts of the codebase. The project underscores the importance of combining manual expertise with automated tools to achieve sustainable performance enhancements at scale.
Sep 18, 2025 3,896 words in the original blog post.
The text introduces the AWS Parallel Computing Service (AWS PCS), a managed service designed to facilitate the running and scaling of high-performance computing (HPC) workloads by utilizing Slurm for scheduling and orchestrating simulations. AWS PCS automates the provisioning and scaling of compute nodes, allowing users to concentrate on refining models rather than managing infrastructure. Despite this, visibility into cluster activity, job performance, and cost drivers remains essential, prompting the integration with Datadog for enhanced monitoring capabilities. This integration provides real-time and historical data insights via Datadog's AWS PCS dashboard, enabling users to optimize HPC workloads, manage costs, and ensure efficient resource allocation. The integration with Slurm offers additional visibility into HPC job activity, assisting in identifying bottlenecks and inefficiencies. Furthermore, by combining these insights with other Datadog integrations, such as Amazon EC2, storage systems, and NVIDIA GPUs, users can quickly determine performance issues and optimize HPC workload performance across different environments.
Sep 17, 2025 1,008 words in the original blog post.
The Windows Certificate Store is essential for maintaining the security and functionality of a Windows environment, as it supports encryption, authentication, and service validation. However, expired or revoked certificates, as well as broken certificate chains, can lead to vulnerabilities and service disruptions. Datadog's Windows Certificate Store integration provides a solution by offering visibility into certificate expirations, broken chains, and Certificate Revocation Lists (CRLs), allowing administrators to proactively address these issues. This integration helps identify and renew expired certificates, set up advance notifications for impending expirations, and monitor CRLs to prevent outages. It also enables the checking of certificate chain validity to ensure uninterrupted service access. By alerting users to potential certificate-related problems, Datadog assists in maintaining the security and performance of Windows services, offering customizable notifications and workflows to fit different environments and needs.
Sep 10, 2025 832 words in the original blog post.
Modern cloud environments are characterized by their complexity and dynamic nature, relying on numerous ephemeral resources, making observability crucial for troubleshooting, maintaining reliability, optimizing performance, and enforcing security standards. As these environments expand, tracking observability coverage becomes increasingly challenging. The Observability overlay in Cloudcraft facilitates the visual identification of critical observability gaps, such as missing Datadog Agents or outdated versions, directly from architecture diagrams. Users can efficiently address these gaps by pivoting to Datadog Fleet Automation, which offers straightforward solutions for installing or updating Agents and ensuring product features are enabled across hosts. This integration between Cloudcraft and Fleet Automation allows users to maintain consistent observability coverage at scale, with both tools being integral to the Datadog platform and available at no additional cost.
Sep 10, 2025 732 words in the original blog post.
Datadog is gearing up for its Tokyo Summit on October 16, 2025, celebrating the community of engineers, SREs, and developers that contribute to its ecosystem. The event will kick off with a keynote by Chief Product Officer Yanbing Li, who will unveil the latest product updates, and feature customer stories illustrating how Datadog aids in scaling and problem-solving. The afternoon will include breakout sessions and workshops, such as the development of LLM-based services for increased reliability, alongside opportunities for hands-on learning. Networking possibilities abound at the expo area, which will feature product demos, AWS GameDay competitions, and a booth from the Japan User Group. An exclusive Partner Day is also planned for the day before the summit, focusing on product updates, sales strategies, and featuring a Capture the Flag challenge. With limited space available, attendees are encouraged to RSVP promptly.
Sep 09, 2025 342 words in the original blog post.
Error handling in Go differs significantly from languages like Java, Python, JavaScript, or Ruby, as it does not automatically generate stack traces when exceptions occur. Instead, Go treats errors as explicit return values, enabling developers to handle them through idiomatic patterns. Over time, Go's error handling capabilities have evolved to include advanced features such as error wrapping, custom error types, and functions like errors.Is, errors.As, and errors.Join, which facilitate error inspection and classification. Despite the lack of built-in error tracing, tools like Datadog's Error Tracking and Orchestrion provide developers with the ability to trace errors effectively by integrating tracing hooks at compile time, allowing for comprehensive visibility into where and how errors occur in Go applications. These capabilities enable efficient debugging and help maintain production-ready, reliable services by clustering related errors, filtering noise, and providing detailed context for each error.
Sep 08, 2025 3,728 words in the original blog post.
The Data Build Tool (dbt) is an open-source analytics engineering framework that helps transform raw data in data warehouses like Snowflake, BigQuery, Redshift, or Databricks using SQL-based workflows. It comes in two forms: the free dbt Core CLI tool and the managed dbt Cloud platform, which offers additional features such as scheduling, UI support, and collaboration tools. dbt introduces software engineering best practices into analytics workflows, including version control, automated testing, data lineage tracking, and CI/CD. It enables data teams to build scalable, trustworthy, and auditable data pipelines by allowing them to write modular SQL transformations with built-in testing for data quality and automated dependency tracking. dbt supports best practices for structuring projects with a layered approach, typically divided into staging, intermediate, and marts layers, to ensure clarity, maintainability, and scalability. Additionally, dbt integrates with CI/CD platforms and monitoring tools like OpenLineage and Datadog to enhance pipeline visibility and reliability.
Sep 05, 2025 1,647 words in the original blog post.
Cloudcraft provides DevOps and security teams with a comprehensive tool for visualizing and managing cloud infrastructure risks, including misconfigurations, vulnerabilities, identity threats, and sensitive data risks. By integrating real-time infrastructure diagrams, Cloudcraft allows teams to quickly identify and prioritize security issues, offering detailed insights and suggested remediation steps. It works alongside Datadog Cloud Security, which continuously scans for misconfigurations and vulnerabilities, ensuring compliance with industry standards. The platform also highlights identity and access management (IAM) risks and locates sensitive data stored within cloud resources, enhancing an organization's security posture. Through Cloudcraft, teams can effectively visualize and address potential security issues, thereby transforming their approach to cloud risk management.
Sep 04, 2025 848 words in the original blog post.
In a guest blog, Suraj Tikoo, an Accenture consultant and Datadog Ambassador, discusses the challenges and solutions involved in implementing a robust testing framework using Datadog Synthetic Monitoring for a client handling sensitive data. The client, already using Datadog for performance monitoring, faced difficulties in adopting flexible testing tools due to strict data guardrails. Tikoo leveraged Datadog's features such as configuration management, reusable modules, test execution scheduling, and reporting to address these challenges, enabling efficient and scalable test suite management. By employing techniques like JavaScript assertions, subtests, and controlled batching, the team was able to integrate complex functionalities like multi-factor authentication and manage limited user licenses effectively, while Terraform was used to ensure test suite backup and recovery. These strategies not only streamlined the client's workflows but also enhanced their testing efficiency by minimizing context switching, enabling automation, and ensuring comprehensive test coverage.
Sep 03, 2025 1,794 words in the original blog post.
In the August episode of This Month in Datadog, Jeremy and Danny discuss several new features and tools aimed at enhancing cloud cost management, Kubernetes infrastructure security, and LiteLLM-powered application monitoring. Key highlights include Datadog Kubernetes Autoscaling, which optimizes costs without compromising performance, and Cloud Cost Management features that provide insights into AWS cost anomalies and enable budget tracking. The episode also introduces Datadog Workload Protection for monitoring user activity in Kubernetes, alongside a LiteLLM integration for improved application visibility. Additionally, upcoming Datadog Summits are announced, offering opportunities for skill development and networking in areas such as AI and security. The episode concludes with updates on collaboration and monitoring tools, encouraging viewers to explore these new features through Datadog's platform and resources.
Sep 03, 2025 539 words in the original blog post.