April 2025 Summaries
34 posts from Datadog
Filter
Month:
Year:
Post Summaries
Back to Blog
Datadog's Snapshot Changes feature offers visibility into all relevant changes in modern multi-cloud environments, providing responders with the information they need to quickly identify and resolve root causes of incidents. By surfacing configuration updates, resource provisioning, and deployment activity directly within existing observability workflows, teams can reduce the time spent investigating incidents and improve overall incident response efficiency. Snapshot Changes also allows users to monitor changes across their multi-cloud infrastructure, view infrastructure changes in context, search by changed fields, and get started monitoring cloud infrastructure changes with one-click access from Monitor Status pages.
Apr 30, 2025
843 words in the original blog post.
Datadog has expanded its Security Orchestration, Automation, and Response (SOAR) solution to bring security automation directly into Datadog Cloud SIEM. Prebuilt, customizable blueprints enable teams to automate key security workflows, enriching, triaging, escalating, and responding to threats without manual effort. Integrated case management streamlines collaboration, while out-of-the-box blueprints help standardize responses to common threats like unauthorized access or malware detection. New SOAR workflows include Identity and Access Management (IAM) workflows that automate responses to suspicious logins and account compromises, Endpoint Detection and Response (EDR) workflows that speed up the investigation and containment of endpoint threats, and Threat Intelligence Enrichment workflows that enrich alerts with external data so teams can prioritize and respond more effectively. Each SOAR blueprint is fully customizable, allowing teams to tailor automation to their environment by modifying steps or conditions to match their processes. Automation reduces response times, improves incident handling, and enables teams to focus on stopping real attacks.
Apr 30, 2025
1,541 words in the original blog post.
Datadog Network Monitoring provides deep visibility into software-defined networking (SDN) solutions such as SD-WANs, cloud-managed networks like Meraki, and data center fabrics like Cisco ACI. With integrations with leading SD-WAN vendors, Datadog offers out-of-the-box dashboards that provide critical device metrics, enabling teams to quickly spot issues and begin troubleshooting. The platform also provides visibility into both physical and virtual components of SD-WANs, allowing for a unified view of how physical connectivity impacts application traffic. Additionally, Datadog's network path feature and enriched NetFlow monitoring add context to the full end-to-end path of SD-WAN traffic and traffic maps from LAN to WAN. The platform also integrates with cloud-managed networks like Meraki, providing actionable insights into Wi-Fi performance issues and bandwidth bottlenecks. Furthermore, Datadog provides end-to-end visibility into Cisco ACI solutions, simplifying data center operations by using policy-driven controls across physical and virtual environments.
Apr 29, 2025
820 words in the original blog post.
Datadog's App and API Protection (AAP) provides detection and defense capabilities to mitigate account takeover (ATO) attacks, which can compromise sensitive information and perform privileged actions. To instrument applications for ATO detection, Datadog automatically instruments supported frameworks like Flask or Node.js, while manual instrumentation requires adding user identifiers and logic for determining legitimate user behavior. Datadog detects ATO attacks by monitoring login activity and flagging suspicious behaviors using built-in detection rules, which provide contextual information to help teams prioritize their response. The platform also provides remediation actions, such as blocking malicious IPs, customizing WAF rules, and creating custom 403 block pages, to slow down or disrupt attacks. To prevent future ATO attacks, teams need to assess the scope of the attack, identify compromised accounts, and adjust detection and response strategies through post-incident analysis. By using Datadog AAP, organizations can protect login endpoints, stop account takeovers before they escalate, and equips their teams with real-time detection and flexible workflows to adapt quickly when attacks do occur.
Apr 28, 2025
1,233 words in the original blog post.
Datastore from Datadog provides fully managed and consistent data storage for automations, allowing teams to create more efficient processes and respond to issues faster. It integrates with the rest of the Datadog ecosystem, enabling fast and persistent storage without external databases. Datastore offers an easy-to-navigate UI, allowing users to view and update data regardless of their technical background. The platform also provides dynamic storage options and read-after-write consistency, making maintenance simple and reducing context switching. With Datastore, teams can dynamically manage data throughout Datadog, maintain a shared data layer from workflows, create full-stack apps with data persistence, and securely store sensitive information. By integrating with App Builder and Workflow Automation, users can build complex automations using their data without leaving the Datadog platform, minimizing context switching.
Apr 25, 2025
1,031 words in the original blog post.
Datadog has reported that 80% of organizations using AWS infrastructure use at least one Infrastructure-as-Code (IaC) tool, but still manually provision infrastructure in production, creating inconsistencies and security risks. To address these challenges, it's essential to identify and close common security gaps in IaC by implementing goals such as inventorying infrastructure with tags, defining IaC and Policy-as-Code rules, shifting policies left into the CI/CD pipeline, and enabling continuous monitoring on your infrastructure. These steps can strengthen security posture and reduce risks from misconfigurations and drift, even in complex environments where ClickOps is prevalent.
Apr 24, 2025
1,426 words in the original blog post.
Kubernetes v1.33 introduces several enhancements aimed at improving security, scalability, and performance. Key updates include Dynamic Resource Allocation (DRA) and the scheduling of extended resources, Container Storage Interface (CSI) improvements, and updates to list API calls that make them less costly. The feature also expands device taints and tolerations, enabling cluster administrators to gain privileged access to devices already in use by other users. Additionally, Kubernetes v1.33 introduces a built-in admission plugin to copy standard node topology labels to pods, promoting a more secure and consistent solution. These updates will enhance storage performance, reliability, and flexibility, making it easier for users to dynamically provision, manage, and scale persistent volumes across diverse environments and storage backends.
Apr 24, 2025
1,197 words in the original blog post.
Datadog Incident Management automates the creation of postmortems by drafting reports automatically from incident timeline data, eliminating manual labor and providing valuable insights. This process speeds up the creation of comprehensive, accurate postmortems that include rich incident data such as monitor events, changes in severity, and communication among responders. The platform also enables collaboration to create actionable knowledge, provides deep context with interactive Notebooks, and facilitates tagging for easy discovery of relevant information. By automating postmortem creation, Datadog empowers teams to save time, learn quickly, and drive meaningful improvements in system reliability and incident response.
Apr 24, 2025
907 words in the original blog post.
Evaluating the functional performance of Large Language Models (LLMs) is crucial in ensuring they continue to work well over time, amid changing trends in production environments. However, producing effective metrics for evaluating LLMs poses significant challenges due to the difficulty in obtaining a stable ground truth and tailoring evaluations to specific use cases. To address this, various evaluation approaches can be considered, including code-based, LLM-as-a-judge, and human-in-the-loop methods. These approaches help characterize the application's performance across different dimensions such as accuracy, relevancy, coherence, toxicity, and sentiment in inputs and outputs. Specifically, context-specific evaluations assess the model's ability to retrieve relevant context and infer from it appropriately, while needle-in-the-haystack tests evaluate the model's retrieval capabilities, and faithfulness evaluations test the model's self-consistency within an LLM-as-a-judge framework. User experience evaluations leverage user feedback data to measure the effectiveness of responses, topic relevancy evaluates the relevance of questions or answers to the application's established domain, and security and safety evaluations monitor for breaches and toxicity in inputs and outputs. By creating a comprehensive monitoring framework, teams can obtain continuous visibility into their LLM application's functional performance and optimize its parameters to improve accuracy, coherence, and user experience.
Apr 24, 2025
2,273 words in the original blog post.
Temporal Cloud is a managed service that enables users to quickly scale the Temporal workflow orchestration engine across their organization, allowing them to focus on developing workflows that increase application reliability. Datadog's Temporal Cloud integration provides granular insights into Temporal Cloud Service, task polling, Workflow activity, and more, enabling users to identify errors and bottlenecks that risk slowing down applications. The dashboard offers metrics such as service latency, task polling rates, and Workflow end states, allowing users to visualize the health of their Temporal Frontend Services, monitor task polling efficiency, and quickly identify errors in their Workflows. By using Datadog's preconfigured Temporal Cloud dashboard, users can gain visibility into their Temporal Workers' task polling, Workflow activity, and more, helping them catch issues such as service latency spikes, failed Workflows, and inefficient task polling.
Apr 24, 2025
1,116 words in the original blog post.
Metaplane, a data observability platform, has joined forces with Datadog to provide machine learning-powered monitoring and column-level lineage, enabling data teams to gain end-to-end visibility into their data ecosystem. This acquisition accelerates Datadog's expansion into data observability, building on its recent launches of Data Jobs Monitoring and Data Streams Monitoring. By unifying observability across applications and data, the combined entity aims to help organizations build reliable data and AI products while bridging the gap between software and data teams.
Apr 23, 2025
239 words in the original blog post.
Rory McCune and Seth Art from Datadog released the 2025 State of DevSecOps study, analyzing tens of thousands of applications and container images across thousands of cloud environments to reveal trends in security posture and best practices. The study found that exploitable vulnerabilities are prevalent in web applications, particularly those using Java; attackers continue to target the software supply chain; usage of long-lived credentials in CI/CD pipelines is still too high but decreasing; only a fraction of critical vulnerabilities are truly worth prioritizing; keeping libraries up to date is a major challenge for developers; minimal container images improve security posture; infrastructure-as-code usage is high in AWS, while ClickOps is still used by many teams. To address these findings, Datadog outlines several best practices, including prioritizing vulnerabilities with runtime context, deploying guardrails within the software supply chain, deploying frequently to stay current on patches, adopting minimal container images, and expanding IaC usage and rein in ClickOps. Additionally, Datadog provides tools like Security Inbox, Cloud Security, Workload Protection, and Code Security to improve security posture and defend against threats.
Apr 23, 2025
2,958 words in the original blog post.
NVIDIA NeMo Evaluator is a microservice with an easy-to-use API that simplifies the end-to-end evaluation of generative AI applications, including retrieval-augmented generation (RAG) and agentic AI. It supports evaluation for a wide range of custom tasks and domains, including reasoning, coding, retrieval, and instruction-following, and allows developers to automatically evaluate their models against academic benchmarks or custom datasets, or score them with standard metrics such as accuracy, ROUGE, BLEU, or LLM-as-a-judge scoring. Datadog LLM Observability can be integrated with NeMo Evaluator to provide end-to-end visibility into the health and performance of LLM applications, tracing requests across RAG components and model inference and evaluation steps, collecting and visualizing key model metrics and metadata, and linking model quality metrics directly to corresponding LLM request traces for unified analysis.
Apr 23, 2025
582 words in the original blog post.
The text discusses the challenges of managing thousands of metrics in high-scale distributed environments and the benefits of using AI-assisted monitoring tools to address these challenges. Tools such as anomaly detection, predictive correlations, and root cause analysis (RCA) automation help by proactively identifying potential problems and providing clearer insights, thereby enhancing system health analysis and speeding up incident response times. The text highlights how these tools, specifically within Datadog's platform, allow users to manage metrics more efficiently by accounting for seasonal patterns, quickly identifying root causes, and understanding the full scope of issues across a distributed system. By utilizing AI-driven features like anomaly monitors, Watchdog Explains, and Metric Correlations, users can swiftly detect, analyze, and resolve critical issues, ensuring smoother operation and faster remediation processes.
Apr 22, 2025
1,163 words in the original blog post.
Synthetic monitoring is a key element of effective UX smoke testing by helping analyze user journeys across deployments via change-resilient automated tests, addressing challenges such as recreating real user behavior, scaling tests, and automating testing workflows. It enables teams to design effective synthetic smoke tests, analyze results, increase efficiency over time, measure the impact and success of their UX smoke testing process, and optimize their coverage by setting up monitor alerts and using dashboards that bring visibility to key metrics. By making synthetic monitoring a core element of your UX smoke testing process, you can increase accuracy, decrease deployment time, and continually improve your app's value by helping it meet—and exceed—user expectations.
Apr 18, 2025
3,880 words in the original blog post.
The text discusses the challenges of managing high-volume logs in distributed systems and the importance of optimizing log management to maintain critical visibility while controlling costs. It highlights the need for organizations to understand how their logs align with their business priorities, reduce noisy log data at the edge, route logs proactively and selectively, and fine-tune log storage on a per-use-case basis. The text also introduces Datadog Observability Pipelines as a solution that can help teams manage their log volumes, generate metrics from logs, and impose rule-based quotas to control log volumes. By implementing these strategies, organizations can make informed decisions on which logs to collect, how they should be handled, and how to manage logging costs effectively.
Apr 17, 2025
1,783 words in the original blog post.
Datadog has been named a Leader in the Forrester Wave: AIOps Platforms, Q2 2025, reflecting its commitment to offering an AI-driven platform that enables customers to observe and secure systems, orient teams, and take action in one place. Datadog processes trillions of telemetry data points every hour and enriches this with unstructured semantic data from various sources for vital context into systems. The platform provides a comprehensive understanding of what's happening in systems, why it's happening, and how DevOps and business teams can solve problems together. With its generative AI engine, agentic AI, event correlation, automation, and unified monitoring features, Datadog enables organizations to accelerate decision-making and operations to improve performance, reliability, and security. By consolidating tools with a single platform, Datadog has helped customers like Tecsys reduce mean time to resolution, spend less time on normalization and centralization of third-party data, and transform their operations by simplifying the work of Site Reliability Engineers and reducing alert incidents by 69 percent.
Apr 15, 2025
549 words in the original blog post.
Nicholas Thomson and Edith Méndez discuss the challenges of managing growing log volumes and the need for modern, cloud-native log management strategies. Organizations are shifting away from legacy solutions due to concerns about migration costs, security risks, performance impacts, and incompatible data formats. Modern solutions offer flexible retention policies, real-time analytical tools, role-based access control (RBAC), and adherence to compliance frameworks. These features help teams reduce costs by selectively keeping essential logs queryable while routing others to long-term archiving. Additionally, modern log management options provide tools for faceted search, real-time search, or analytics that require little or no technical experience, improving collaboration across teams and enabling faster issue resolution. The authors also explore the importance of targeted insights for faster incident response and security vulnerability detection. To successfully migrate logs, teams must segment data sources, enforce compliance with sensitive data redaction, and unify log formats through standardization techniques. Finally, Datadog Observability Pipelines provides a cost-efficient and disruption-free migration solution that simplifies log migration and enables teams to configure log collection, transformation, and routing without disrupting existing workflows.
Apr 14, 2025
1,867 words in the original blog post.
Mallory Mooney discusses the importance of risk assessment in cloud environments and how it requires context beyond just monitoring activity. She identifies common categories of risky behavior, including anomalous user and admin activity, identity risks, and resource misconfigurations. To connect these behaviors to specific entities, she highlights the need for entity analytics, which correlates logs with users, service accounts, and roles. Datadog Cloud SIEM Risk Insights provides a comprehensive approach by aggregating security logs, analyzing patterns, and generating alerts based on predefined and custom rules. It also maps events to identities and resources, providing a better understanding of what the risky behavior is and how it should be prioritized. By leveraging these capabilities, organizations can identify and respond to potential threats in their cloud environments.
Apr 14, 2025
1,750 words in the original blog post.
We've faced challenges with frontend monitoring at Datadog, including flaky acceptance tests, tool sprawl, and limited visibility into user behavior. To address these issues, we enhanced our digital experience monitoring (DEM) products by integrating Synthetic Monitoring, RUM's Session Replay, and Error Tracking. This enabled us to catch problems before they affected users, accelerate debugging, and streamline our investigative workflows. By using Synthetic Monitoring, we transitioned from acceptance tests to browser tests, reducing costs and maintenance requirements. Datadog Continuous Testing allows us to run tests in parallel, at scale, and in multiple environments. Session Replay provides maximum visibility into user journeys, while Error Tracking simplifies the troubleshooting process by embedding critical contextual information within each issue. By incorporating these tools, we've optimized our frontend testing and debugging, enabling us to ship high-quality features faster and quickly address issues impacting users.
Apr 11, 2025
1,718 words in the original blog post.
The Continuous Profiler team at Datadog has introduced enhancements to their flame graph visualization, making it more intuitive and accessible to engineers. The new color coding scheme simplifies the interpretation of frames, with darker shades indicating higher resource usage. Additionally, a new call graph visualization summarizes profiling data by highlighting relationships among methods called in a service and their impact on resources. The call graph displays each method as a single node, with edges used to convey which methods have called each other, making it easier to discern method calls and optimize resource usage. These improvements aim to make profiling a standard practice for all developers by reducing the learning curve and providing valuable insights into application performance.
Apr 11, 2025
1,642 words in the original blog post.
Self-Service Actions in Datadog Software Catalog enable developers to act independently while ensuring alignment with organizational standards for security, compliance, and reliability. This allows platform teams to create standardized templates through the App Builder, reducing wait times for approvals and manual, repetitive tasks that could be automated. With Self-Service Actions, developers can easily provision cloud infrastructure, manage infrastructure with increased efficiency and security, accelerate remediation and incident resolution, scaffold new services with embedded best practices, and extend and customize existing tooling. These actions reduce misconfigurations, inconsistencies, and security gaps by providing a single platform for developers to access tools and functionality directly within Datadog.
Apr 10, 2025
909 words in the original blog post.
Self-Service Actions in Datadog's Internal Developer Portal aim to streamline operations for engineering teams by empowering developers to independently manage infrastructure and resolve incidents while adhering to organizational standards. This feature addresses the common bottlenecks caused by developers' reliance on platform engineers for routine tasks, which often lead to delays and inefficiencies. By offering pre-approved templates and automated workflows, developers can quickly provision cloud resources, manage infrastructure, and accelerate incident resolution without leaving the Datadog platform. Additionally, the integration allows for scaffolding of new services with embedded best practices, and the ability to extend and customize existing tooling to fit unique organizational needs. Through these capabilities, Self-Service Actions enhance efficiency while maintaining security and compliance, allowing teams to work faster and more effectively.
Apr 10, 2025
1,009 words in the original blog post.
The Datadog team reengineered their AWS Lambda extension to deliver high-fidelity telemetry with minimal overhead, resulting in a 82% reduction in cold start latency and a 40% reduction in memory usage. The rewrite was done using Rust, which provided significant benefits including tiny binaries, excellent concurrency primitives, and memory safety enforced at compile time. The team carefully balanced design constraints such as minimizing impact on the running Lambda function handler code and reducing CPU consumption during the invoke phase. They also implemented a failover strategy to deliver performance improvements to customers while having a gradual migration process. The project demonstrates the potential of Rust for Lambda extensions, with reduced binary size from 55 MB to just 7 MB, and extremely fast cold start times of around 70 ms.
Apr 09, 2025
2,022 words in the original blog post.
Datadog has integrated with Microsoft's core suite of data protection and disaster recovery services, including Azure Backup Vault, Azure Recovery Services Vault, and Azure Site Recovery. This integration enables users to track the status of their vaults' backup jobs and measure their organization's data loss following disaster recovery using recovery point objectives (RPO). The integration provides insights into backup vaults from both Azure Backup Vault and Azure Recovery Services Vault, as well as health status updates from recent backup jobs. It also allows users to monitor their recently run backup jobs and gain quick insights into their health status or whether they failed. Additionally, Datadog's monitor template for Azure Backup Vault job errors enables users to set up alerting for errors, ensuring that engineers can investigate and restore the health of their backup jobs promptly. The integration helps ensure that Azure environment backups stay healthy and that users are prepared when disasters occur, making it an essential tool for organizations relying on these services.
Apr 08, 2025
746 words in the original blog post.
Oracle NetSuite is a fully managed business management platform that centralizes and automates core business functions. However, its custom implementations using SuiteScript can be challenging to monitor and optimize. Datadog has partnered with Continuous AI to develop an integration that enables visibility into performance, security, and access patterns, allowing users to track script health, identify potential degradation, and measure the impact of script failures on business operations. The integration provides real-time alerts for script issues, error patterns, or performance issues, enabling users to quickly investigate and resolve problems before they affect critical business operations. Additionally, it offers visibility into user access patterns, system changes, and security events, allowing users to secure their NetSuite environment and promote compliance with regulatory requirements.
Apr 07, 2025
995 words in the original blog post.
Security and SRE: How Datadog's combined approach aims to tackle security and reliability challenges
At Datadog, the company is proactively changing the industry paradigm by integrating Security and Reliability Engineering (SRE) principles to address security challenges. By merging SRE and security groups into a single organization, they've unified all aspects of their operational and security posture, enabling practical vetted SRE solutions to be applied to security challenges and vice versa. This approach has led to enhanced incident response, improved risk management, governance, and system reliability, as well as strengthened security culture by breaking down silos between development, operational, and security teams. The combined approach has also enabled security teams to participate in many types of incidents, which gives them hands-on experience to efficiently respond to security breaches. Additionally, it has made an impact in several areas of their work, such as improving log governance, auditing, and security control rollout, and enhancing security response by creating a singular incident process. The organization is structured around three key areas: product, internal cloud infrastructure, and operations, each with teams focused on security to ensure the Datadog platform is built with customers' trust and safety in mind.
Apr 07, 2025
971 words in the original blog post.
Datadog Real User Monitoring (RUM) has introduced an Optimization page that provides deep insights into web performance, helping teams pinpoint the root cause of browser performance issues. The new feature uses real traffic data to offer information about Core Web Vitals, custom Loading Time metric, and recurring errors across different user segments. It enables users to quickly identify and resolve user experience issues, optimize load speed with LCP insights, minimize input lag and optimize responsiveness with INP insights, and improve visual stability with CLS insights. The Optimization page provides a unified view of web performance, equipping teams with the tools needed to enhance user experiences and deliver faster, more engaging applications.
Apr 04, 2025
973 words in the original blog post.
Cloudflare is a content delivery network (CDN) that helps businesses accelerate, protect, and optimize their websites, applications, and APIs. Cloudflare logs provide detailed insights into HTTP requests, including origin and response metadata, as well as security, TLS, and encryption information. The logs are structured as JSON objects, with each entry representing a single HTTP request processed by Cloudflare. Key fields in the log include EdgeStartTimestamp, EdgeEndTimestamp, ClientRequestQuery, EdgeResponseStatus, CacheStatus, OriginIP, OriginTLSVersion, OriginResponseDurationMs, WAFAction, BotScore, and ThreatScore. These fields can be used to monitor Cloudflare logs for debugging and troubleshooting, managing cost, security monitoring and threat detection, compliance and auditing, and monitoring in a unified platform with Datadog.
Apr 03, 2025
2,037 words in the original blog post.
In a rapidly evolving digital environment, ensuring seamless user experiences is paramount for organizations aiming to maintain their reputation and revenue. Digital experience monitoring (DEM) has gained prominence as a method for optimizing user interactions with business-critical applications. Datadog's Synthetic Monitoring tool aids in this endeavor by offering end-to-end testing for websites, mobile applications, and APIs, allowing issues to be identified and resolved before affecting users. Recent updates to the platform enhance the test creation process with features like synthetic test templates, browser test recommendations, revamped multistep API test creation, and an Element Inspector for mobile app testing. These updates facilitate quicker DEM adoption by simplifying test setup, improving accuracy, and ensuring tests align with real user behaviors, thus enhancing the overall reliability and performance of applications across various devices.
Apr 03, 2025
1,041 words in the original blog post.
AWS customers can now securely and cost-effectively connect to Datadog services from any AWS region without the need for complex networking configurations. Cross-region PrivateLink simplifies management, enhances security, and reduces costs compared to traditional solutions like VPC peering, NAT gateways, and AWS Transit Gateways. This approach also makes it easier to confidently monitor multi-region workloads while adhering to data privacy standards, especially in regulated industries such as healthcare and finance. By using cross-region PrivateLink, organizations can reduce costs across networking, storage, and ingestion while maintaining full control and visibility into their observability data. To set up AWS cross-region PrivateLink for Datadog, customers need to follow simple steps, including creating a PrivateLink VPC endpoint, configuring the interface endpoint, and testing the connection.
Apr 02, 2025
800 words in the original blog post.
PCI DSS compliance is critical for organizations that handle payment card transactions to protect customer data and maintain trust. The Payment Card Industry Data Security Standard (PCI DSS) was created to establish best practices for securing cardholder data, but meeting its requirements can be complex, especially with evolving compliance standards. PCI DSS v4.0.1 introduces minor modifications compared to previous versions, prioritizing overall security posture over specific tools and technologies. To achieve PCI DSS compliance, organizations can use Datadog's security solutions, which automate vulnerability detection, enhance threat management, and provide continuous monitoring. By integrating these measures, businesses can strengthen their cybersecurity posture and improve overall resilience while meeting PCI DSS requirements.
Apr 01, 2025
1,597 words in the original blog post.
The FinOps Open Cost and Usage Specification (FOCUS) is an open standard that simplifies the process of understanding and managing cloud costs by normalizing cost and usage data from multiple providers. It was created by a group of FinOps practitioners, CSPs, SaaS providers, and other contributors in an open source project led by the FinOps Foundation. FOCUS provides a vendor-agnostic format for providers to deliver their cost and usage data, enabling users to analyze costs from all providers. The specification defines a set of columns that describe charges, including fields such as charge description, billed cost, service name, consumed quantity, and consumed unit. It also includes features like normalization, flexibility, and maintenance by the FinOps Foundation. FOCUS has been adopted by major public clouds and is being actively used by Datadog to provide a unified view of cloud costs. By adopting FOCUS, organizations can gain complete visibility into their total cloud spend, launch and mature a FinOps practice, and make data-driven decisions to optimize their cloud spending.
Apr 01, 2025
1,731 words in the original blog post.
This Month in Datadog is a monthly update of the company's latest features, product announcements, and more. The March episode covers several new features, including Attacker Clustering, which identifies and groups attacker behaviors during distributed attacks, Auto Test Retries, which automatically retries failing tests up to five times, and new Observability Pipelines integrations with Amazon S3, Amazon Data Firehose, and AWS Lambda, as well as SentinelOne. Additionally, Reference Tables is now generally available, enabling teams to upload custom metadata to enrich their Datadog telemetry with business-critical context. The episode also features blog posts on creating an effective paging strategy and structuring on-call rotations, as well as a quick look at upcoming events and webinars.
Apr 01, 2025
610 words in the original blog post.