Home / Companies / Datadog / Blog / March 2025

March 2025 Summaries

22 posts from Datadog

Filter
Month: Year:
Post Summaries Back to Blog
With Reference Tables in Datadog, users can import and enrich their logs, events, and other telemetry data by adding business-critical information from various sources such as CSV files, cloud storage, or SaaS integrations. This enables faster troubleshooting, more efficient business processes, and lower costs by providing meaningful context to the data. Reference Tables allow users to define new entities like customer details, service names, and IP addresses, which can be used across various business, application, and security use cases. By enriching logs with dynamic metadata, users can correlate and analyze data in real-time, making it easier to identify security threats, troubleshoot performance issues, and optimize cloud costs. Additionally, Reference Tables enable advanced querying capabilities using tools like Sheets, DDSQL Editor, and Log Workspaces, allowing users to perform complex analysis and build reports without requiring technical expertise.
Mar 31, 2025 966 words in the original blog post.
Datadog has introduced private actions, which enable users to build workflows and apps to securely manage their self-hosted infrastructure. These private actions support over 300 actions across six connection types, including Kubernetes, GitLab, Jenkins, PostgreSQL, Temporal, and HTTP. With private actions, teams can automate remediation in Kubernetes by using any of the 150+ supported Kubernetes actions, such as automatically restarting deployments to reduce downtime or creating an app that lets them manage deployments, pods, and containers directly from Datadog. Private actions also provide a unified interface for visibility and action, allowing teams to consolidate mission-critical information about their deployments, pods, and containers without requiring Kubernetes expertise or logging in to the Kubernetes console.
Mar 31, 2025 990 words in the original blog post.
The Azure Service Bus integration with Datadog provides deep visibility into message queues and topics, surfacing granular metrics for each of the subscriptions in each service bus topic. This allows teams to troubleshoot delays, failures, and bottlenecks more efficiently, identify potential issues early, scale resources as needed, and resolve bottlenecks before they affect customer experience. The integration helps teams proactively monitor message queues and topics, ensuring smooth communication between services and preventing delays that could impact customers. It also provides real-time metrics for key Azure Service Bus metrics such as active messages, dead-lettered messages, throttled requests, and queue space usage, enabling teams to optimize message flow, detect processing issues, and take proactive actions before they impact application performance.
Mar 31, 2025 968 words in the original blog post.
Datadog has introduced Process Tag Rules, which enable users to enrich their system's end-to-end visibility by deriving tags from a process's command line and creating performance metrics and alerts specific to those tags. This feature allows users to quickly find common issues with groups of processes, track system health, and create custom process groups for more effective monitoring. By using Process Tag Rules, users can gain deeper insights into their infrastructure and services, troubleshoot issues more efficiently, and optimize resource allocation. The feature is available to Datadog customers, including a 14-day free trial for new users.
Mar 27, 2025 911 words in the original blog post.
Datadog teams create on-call rotations to ensure continuous uptime for their critical services. The size of the team largely determines the structure of a rotation, balancing service coverage with a sustainable workload for responders. Team size shapes how the rotation works and can affect the experience of everyone in the rotation. Small teams often use 24/7 rotations with brief turns in the on-call role and fewer teammates available to alternate. The shift length varies, but is generally between eight and 12 hours, aiming to maximize effectiveness while minimizing fatigue. Engineers should do only on-call work as much as possible during their on-call days, separating feature work from on-call responsibilities. Responders receive comprehensive support before, during, and after their on-call duties, including training, resources, and backup secondary responders. Managers participate in their teams' rotations to understand procedures and improve the experience. Datadog provides tools and platforms to facilitate effective on-call practice, such as On-Call, which integrates monitoring, paging, and incident response into a single platform.
Mar 26, 2025 1,749 words in the original blog post.
To create an effective on-call process, organizations should focus on responder attention on the most important issues and help facilitate a sense of ownership over them. This can be achieved by designing symptom-based pages and monitors that detect user impact, establishing clear lines of responsibility for on-call team members, and unifying paging, remediation, and analysis across platforms using tools like Datadog Incident Response. By refining their alerts to highlight problems with the most customer impact and integrating features that help track the entire lifecycle of an incident, organizations can empower engineers and keep their focus on high-impact work, ultimately reducing burnout and costs associated with inefficient on-call and incident management processes.
Mar 26, 2025 1,916 words in the original blog post.
Datadog Observability Pipelines now integrates with Google SecOps, a cloud-native SIEM solution, to help organizations manage their security data by centralizing log collection and extract, transform, and load (ETL) processes within their own infrastructure. With this integration, users can standardize log collection and processing before routing logs to Google SecOps, enriching logs with GeoIP information and redacting sensitive data, and normalizing security logs from various sources using the Open Cybersecurity Schema Framework (OCSF). The integration also enables users to send logs directly to Google SecOps for AI-powered threat detection and automated response playbooks. Additionally, Observability Pipelines provides flexibility in routing logs to multiple destinations, including SIEMs, data lakes, logging platforms, and cloud storage providers, allowing organizations to test new tools without disrupting existing workflows.
Mar 25, 2025 653 words in the original blog post.
In the blog post, Bowen Chen addresses the growing complexity of maintaining compliance and minimizing security risks in cloud-based and AI-driven environments, emphasizing the importance of adopting a shift-left approach to proactively address issues early in the development lifecycle. The article highlights how traditional security tools focus on runtime detection, which can leave organizations vulnerable until issues are discovered. By combining shift-left practices with runtime solutions, organizations can better prepare for regulatory compliance by implementing measures such as redacting sensitive data in non-production environments, identifying infrastructure misconfigurations before deployment, and analyzing third-party dependencies for vulnerabilities. The use of tools like Datadog for static code analysis, infrastructure-as-code scanning, and policy as code is suggested to create built-in safety checks and help prevent misconfigurations and vulnerabilities from reaching production. This multi-layered approach aims to enhance security and compliance postures, ensuring sensitive data remains protected and organizations avoid potential breaches and financial repercussions.
Mar 25, 2025 2,366 words in the original blog post.
HTTP headers play a crucial role in web app network communication, providing specifications for activities such as data handling and session verification. However, insecure HTTP headers can be exploited by attackers to breach apps in various ways, including cross-site scripting (XSS), web-cache poisoning, clickjacking, and man-in-the-middle (MITM) attacks. To combat these threats, configuring security-focused HTTP header fields is essential, which can be challenging due to the variety of data they contain. Synthetic testing enables developers to check their security header configuration and spot potential weak points in their app, better securing existing headers and configuring new ones as necessary. By using synthetic testing tools like Datadog Synthetic Monitoring, developers can ensure that their security headers are implemented correctly and aren't exposing key information or entry points for attackers, ultimately protecting their apps against various types of attacks.
Mar 24, 2025 1,291 words in the original blog post.
State, local, and education (SLED) organizations often struggle with managing log data due to scattered and fragmented systems, leading to inefficiencies and increased security risks. The use of Datadog Observability Pipelines is proposed as a solution to centralize and streamline log management by enabling the ingestion, enrichment, and deduplication of logs without discarding existing tools. This approach allows for dual shipping of logs, giving organizations the flexibility to maintain their current systems while also leveraging Datadog's capabilities for better log visualization, analysis, and collaboration. Additionally, by reducing log noise and managing costs through filtering, SLED organizations can improve efficiency and security without exceeding tight budget constraints. The centralized processing of logs through Observability Pipelines facilitates cross-team collaboration and enhances investigative workflows, ultimately improving the visibility and effectiveness of public services while maintaining compliance with security standards.
Mar 19, 2025 1,133 words in the original blog post.
The text discusses the importance of balancing security and compliance in today's cloud landscape. It highlights the challenges of integrating these two goals and provides guidance on how to measure an organization's security posture using various metrics, including mean time to detect (MTTD), mean time to acknowledge (MTTA), mean time to resolve (MTTR), intrusion attempts, false positive rates (FPR), security incidents, governance effectiveness, compliance, and preparedness. The text also touches on the use of Service Level Objectives (SLOs) as a framework for gauging success and connecting security to broader operational goals. It emphasizes the need for organizations to tailor their approaches to meet their unique goals and provides examples of how Datadog can help bridge the gap between security and compliance by enabling real-time monitoring and analysis of these key metrics.
Mar 18, 2025 2,441 words in the original blog post.
Datadog provides visibility into an organization's security posture across three key areas: response and remediation, incidents and threats, and governance, compliance, and preparedness. Datadog Cloud SIEM automatically tracks mean time to detect (MTTD), mean time to acknowledge (MTTA), and mean time to resolve (MTTR), providing a built-in overview dashboard for reviewing each metric alongside other data such as signal trends. Datadog Flex Logs decouples the cost of log storage from querying, providing short- and long-term log retention without sacrificing visibility. The platform also provides tracking SLOs, false positive rate (FPR) metrics, and security incident data to help organizations measure their level of preparedness and overall compliance. Additionally, Datadog offers built-in compliance reports for quickly identifying gaps in the environment, automated security baselines, and scorecards to simplify the process of applying checks across services and monitoring their status. By providing this visibility into key areas of an organization's security posture, Datadog enables teams to continually monitor their services for security issues, minimize costly risks, and prioritize security-focused metrics, goals, and events that matter most to the organization.
Mar 18, 2025 1,974 words in the original blog post.
Datadog organizations are logical groups of users, configurations, and telemetry data that have a parent-child relationship. Large enterprises need to fulfill legal requirements and ensure isolation among divisions by implementing a multi-organization setup in Datadog. This setup provides benefits such as consolidated billing data, logically isolated data, centrally governed configurations, user flexibility, and the flexibility to have different structural organization models. The recommended approach is to analyze organizational requirements with technical account managers or customer success managers, keep child organizations low, create Restricted Datasets for limited observability sharing, use parent organizations only for central management and billing, automate provisioning of users and child organizations, standardize integrations with cloud environments, and implement continuous governance through automated processes.
Mar 14, 2025 1,486 words in the original blog post.
Datadog's integration with GitHub Copilot enables organizations to monitor and analyze the usage of AI-powered coding tools in their workflows. By tracking license distribution and user engagement, teams can gain insights into how Copilot is being used and identify areas where it creates the greatest impact. The dashboard provides visibility into metrics such as active seats, inactive seats, suggestions accepted, and average lines per suggestion across various programming languages and IDEs. This information helps teams adjust their cloud spend on licenses and introduce Copilot to new projects and teams that are likely to benefit from its features. By analyzing usage trends and identifying compatible coding languages and teams, organizations can optimize the use of Copilot and make data-driven decisions about its adoption.
Mar 11, 2025 761 words in the original blog post.
Datadog is hosting its first Summit in São Paulo on April 29, celebrating its community of engineers and developers who contribute to the platform's growth. The event will feature a keynote presentation by Datadog Director of Engineering Daniela da Cruz, as well as three hands-on workshops on topics such as observability with logs, OpenTelemetry, and synthetic and real user monitoring. Participants will also have opportunities to network with local users, attend product demos, and compete in an AWS GameDay for prizes. The summit is invite-only, but Datadog hosts events throughout the year, including its annual conference DASH, which registration is now open for.
Mar 11, 2025 290 words in the original blog post.
The MITRE ATT&CK Map is a feature of the Datadog Cloud SIEM that provides security teams with clear visibility into potential threats and helps them proactively defend against cyberattacks. The map visualizes detection coverage across different attack surfaces, allowing analysts to assess their overall coverage, identify gaps, and refine their SIEM strategy. With real-time visibility into enabled rules and their data sources, security teams can streamline rule creation and strengthen their defenses by creating custom rules with pre-populated tactic and technique tags. By using the MITRE ATT&CK Map, security teams can improve detection coverage, prioritize threats, and enhance their overall security posture.
Mar 10, 2025 823 words in the original blog post.
Datadog's Application Security Management (ASM) can help security teams improve the effectiveness of their web application firewalls (WAFs) by addressing their limitations. WAFs struggle with detecting and preventing exploits due to their reliance on patterns, which can lead to false positives and false negatives. Datadog ASM's In-App WAF detects and blocks suspicious requests based on known attacker behavior, providing a more accurate picture of the threat landscape. Additionally, Datadog Exploit Prevention instruments applications at a deep level to identify vulnerability exploits, reducing false positives and false negatives. By extending their capabilities, security teams can make the most of their WAFs' strengths while addressing their weaknesses.
Mar 06, 2025 1,442 words in the original blog post.
The episode of "This Month in Datadog" features an in-depth conversation between two leaders, Chief Product Officer Yanbing Li and Senior Vice President of Engineering David Mitchell, discussing their approach to building products, engineering culture, AI, and more. The conversation also includes coverage of the new platform for the Datadog Certification Program and a sneak peek at DASH 2025, which is returning to New York City in June. New features and updates released this month include Cloud SIEM Risk-based Insights for AWS, Service Catalog now Software Catalog with Schema v3.0, Diagram AWS infrastructure in Datadog with Cloudcraft in Preview, monitoring GPU and data transfer costs in Kubernetes environments, and a new integration to monitor Google Cloud TPUs. The episode concludes by encouraging viewers to check out the release notes and sign up for a 14-day free trial to experience these features and updates firsthand.
Mar 05, 2025 245 words in the original blog post.
Datadog's Cloud Security Management (CSM) Identity Risks offers actionable insights to detect cross-account access risks in AWS, helping organizations identify and remediate misconfigured IAM roles that can lead to privilege escalation and unauthorized access. These risks include IAM roles assuming another role with admin privileges, EC2 instances assuming an IAM role with admin privileges, and IAM users assuming an IAM role with admin privileges. Datadog's CSM Identity Risks continuously scans cloud infrastructure to find security risks, providing visibility into permissions gaps, unused and unnecessary admin privileges, and potential cross-account access. By identifying these risks, organizations can take steps to mitigate them and safeguard their cloud identities and resources from attacks that take advantage of IAM misconfigurations.
Mar 05, 2025 873 words in the original blog post.
Google has introduced Core Web Vitals as a standardized set of metrics to monitor user experience in web applications. These vitals provide developers with data-backed insights to optimize their app's performance, but implementing them can be challenging, especially for single-page applications (SPAs). To help troubleshoot poor Interaction to Next Paint (INP) scores, Datadog provides features such as recording soft navigations and contextualizing INP data with historical information and other Core Web Vitals. By using Datadog's RUM platform, developers can measure interactivity, quickly identify latency issues, and pivot to Product Analytics to analyze user behavior and correlate variations in INP performance with broader trends in user satisfaction.
Mar 04, 2025 1,657 words in the original blog post.
Java remains a widely used programming language, particularly for enterprise backend systems, due to its robust runtime, portability, and extensive ecosystem of libraries. However, developers face challenges when deploying Java applications in cloud environments like Kubernetes, especially with older frameworks, due to Java's higher memory and CPU requirements and slower startup times. To optimize performance, it is crucial to tune Java applications by selecting appropriate JVM versions, frameworks, and server software, configuring Kubernetes deployments, and adjusting memory and CPU limits. Modern solutions such as GraalVM and frameworks like Quarkus can enhance performance and reduce startup times. Additionally, garbage collection tuning and monitoring tools like Datadog can provide insights into application performance, helping to address issues such as GC pauses and OOM kills. By understanding and applying these strategies, developers can improve the efficiency and scalability of Java applications in containerized environments.
Mar 04, 2025 4,130 words in the original blog post.
At Datadog, the Cloud Security team faces numerous challenges in securing its complex infrastructure due to finite resources and time constraints. The team adopts a Find, Fix, Remediate, Prevent (FFRP) methodology to tackle risks effectively and avoid the "security treadmill." This approach helps identify systemic root causes of issues, fix them, remediate downstream effects, prevent similar problems from occurring, and establish guardrails to contain future incidents. The team utilizes Datadog Cloud Security Management (CSM) to implement their FFRP methodology, partnering with engineering teams to carry out cloud security strategy and ensuring the right data is sent to the right stakeholders. By adopting this approach, the Cloud Security team reduces open vulnerabilities in CSM to zero and empowers engineers to remediate cloud security risks themselves, without relying on extensive Cloud Security personnel intervention.
Mar 03, 2025 1,773 words in the original blog post.