Home / Companies / Datadog / Blog / May 2024

May 2024 Summaries

14 posts from Datadog

Filter
Month: Year:
Post Summaries Back to Blog
The Device Topology Map in Datadog Network Device Monitoring (NDM) provides a clear and concise view of an enterprise network's relationships and dependencies, enabling network engineers to quickly troubleshoot issues by visualizing the topology at a glance. The map uses dynamic device links, detailed health metrics, and granular tagging to help identify problematic devices and reduce mean time to resolution (MTTR). By filtering devices based on tags or custom color rules, teams can pinpoint issues across their network and start troubleshooting immediately, with features such as hover-over panels providing additional information about individual devices.
May 29, 2024 831 words in the original blog post.
The 2024 State of DevSecOps study by Christophe Tafani-Dereeper highlights significant findings concerning the security of applications and cloud environments, emphasizing the prevalence of third-party vulnerabilities, particularly in Java services. The study underscores the need for visibility into these vulnerabilities through tools like Datadog Software Composition Analysis (SCA) and stresses the importance of prioritizing remediation efforts using frameworks and standards such as CVSS and EPSS. It also advocates for minimal container images to reduce vulnerabilities and highlights the benefits of infrastructure-as-code for zero-touch production environments. Furthermore, the study recommends the use of short-lived cloud credentials in CI/CD pipelines to mitigate the risks associated with long-lived credentials. Datadog's security platform is presented as a comprehensive solution for identifying and prioritizing security threats, offering tools to improve security posture through effective monitoring and remediation strategies.
May 29, 2024 2,526 words in the original blog post.
We replaced the Java-based static analyzer with a Rust-based one, which tripled our performance and resulted in a tenfold reduction in memory usage. We achieved this by leveraging Tree-sitter's parsing capabilities and GraalVM's JavaScript runtime. The migration was successful despite being the first major Rust project for Datadog, thanks to a firm grasp of Rust concepts within 10 days, strict coding standards, and the use of specialized tools and libraries. The new analyzer provides faster feedback on pull requests and improved performance in resource-constrained environments like CI systems with limited RAM and CPU resources.
May 23, 2024 1,891 words in the original blog post.
Thomas Sobolik, one of Datadog's Engineering VPs, shared his career journey and insights on leadership, technology, and company culture in this edition of the Datadog Engineering Spotlight. Ivo Dimitrov, another Engineering VP, discussed his background as an engineering manager for top organizations, his transition to management, and what excites him about building distributed data systems at Datadog. Sobolik reflected on his 30-year career, starting with a passion for electrical engineering and software development, and transitioning from individual contributor to manager. He highlighted the importance of teamwork, innovation, and continuous learning in his role as a manager. Sobolik discussed his experience working with product teams, helping colleagues, and navigating the challenges of growing a company like Datadog. He emphasized the need for a balance between introducing new processes and respecting the existing substrate to avoid resistance. The interview also touched on the importance of humility, experimentation, and taking intelligent risks in the engineering culture at Datadog.
May 21, 2024 2,647 words in the original blog post.
DORA metrics provide a holistic view of the software development life cycle, offering insights into speed and stability. The four key metrics - deployment frequency, lead time for changes, time to restore services, and change failure rate - bridge the gap between production system-based metrics and development-based ones. To effectively collect and analyze DORA metrics, organizations need to establish a unified process of monitoring, standardize data points, and prioritize insights based on team goals. This can be achieved by defining success criteria for deployments, identifying failures or responses, determining incident start and end times, and selecting appropriate time spans for analysis. By using DORA metrics to identify inefficiencies in CI/CD workflows, teams can adjust their tooling to create more effective processes. Datadog offers a range of tools to help optimize developer workflows based on DORA metrics findings, including performance monitoring, continuous integration, code quality, and incident response features.
May 21, 2024 2,476 words in the original blog post.
Modernizing your CI/CD environment is crucial as you adopt new technologies and scale your workloads, enabling agile and resilient workflows. However, this process can be lengthy and complex, especially when migrating to a new provider. Datadog provides visibility into your entire CI/CD environment, even during modernization, with features such as pipeline monitoring, autoscaling runners, tracing, and continuous delivery tools integration. By adopting modern CI/CD tools and best practices, you can streamline troubleshooting workflows, proactively monitor pipelines, and gain full visibility into your CI/CD stack. When migrating to a new provider, it's essential to use Datadog for post-migration validation, optimization, and continuous improvement, ensuring that your pipeline performance and reliability are maintained.
May 20, 2024 2,022 words in the original blog post.
Datadog's .NET profiler is a comprehensive tool for monitoring application performance. It provides detailed insights into CPU consumption, wall time, exceptions, lock contention, and memory usage profiling. The profiler uses a pull model to collect data, which means it retrieves CPU consumption details from the operating system every minute. Memory usage profiling is useful for identifying high CPU consumption due to excessive garbage collection by the runtime, pinpointing specific parts of code responsible for memory allocation, and viewing samples of objects that stay in memory after garbage collection. The profiler tracks allocations using the `AllocationTick` event, which provides information about the last allocated object crossing a 100 KB threshold. It also monitors the lifetime of sampled allocations using weak handles created from the given address and added into the list of monitored objects with its creation time. However, upscaling issues arise due to complications in scaling sampled values to estimate real values, particularly for live objects. The profiler does not upscale sampled values currently, but it may change in the future when the allocation sampling distribution allows more realistic upscaling.
May 20, 2024 1,944 words in the original blog post.
Google Cloud Next 2024 was attended by over 30,000 people who saw the latest developments in Google Cloud's AI stack, including upgrades to foundation models and new ways to use them. These improvements include Gemini 1.5 Pro, a powerful model with multimodal input support, and new variants of the open source Gemma model. Additionally, Google introduced its new Vertex AI Agent Builder, which enables users to build AI agents via conversational prompts. The event also saw new investments into performant hardware, including Google's most powerful TPU yet, v5p, and Arm-based Axion GPUs. Furthermore, Google highlighted its bolstered cloud security features, including the use of AI to detect and prevent security threats, as well as integrations with key services like Google Threat Intelligence and Security Command Center. The event also saw the introduction of code assistance tools, such as Gemini Code Assist and CodeGemma, designed to improve developer efficiency. Lastly, Google emphasized the importance of monitoring and securing one's Google Cloud environment, with features like CI Pipeline Visibility and Intelligent Test Runner.
May 17, 2024 968 words in the original blog post.
Felix Geisendörfer introduces a tool for continuous profile-guided optimization (PGO) for Go, designed to reduce CPU usage by up to 14% in Go services. The tool, datadog-pgo, integrates with the Go build process by using representative CPU profiles from production environments to optimize machine code, particularly in inlining and devirtualization of function calls. Testing on Datadog's internal services demonstrated significant CPU usage reduction, translating to substantial cost savings, even in already optimized services. Despite challenges such as a discovered memory usage increase due to an inlining issue in a gRPC library, Datadog implemented a mitigation and proposed a new goroutine stack profiler to the Go project. The tool, available for a free trial, promises further optimizations as Go's PGO capabilities evolve, with plans for wider deployment and more detailed results from Datadog's fleet-wide rollout.
May 13, 2024 1,400 words in the original blog post.
The Datadog Slack integration helps teams streamline their incident management by minimizing context-switching and simplifying collaboration. It enables seamless end-to-end documentation of every incident, which can offer enormous advantages when it comes to building resilience. With the `/datadog incident` command, anyone can declare an incident from any Slack channel with the Datadog integration, setting down key information, assigning roles, sending custom notifications, and creating a dedicated Slack channel for the incident. The integration also provides an action tray within each incident Slack channel created with Datadog, allowing responders to quickly update the incident status and description, add responders, page on-call team members, navigate to the incident timeline in the Datadog app, or start a Zoom meeting for channel members with a single click. Additionally, the integration helps document incidents in detail for postmortem analysis by mirroring all messages in incident channels to associated incident timelines in Datadog and providing an action tray within each incident Slack channel created with Datadog.
May 07, 2024 1,199 words in the original blog post.
The Container Image Trends view in Datadog provides insights into every container image used in an environment, helping users quickly detect and remediate security and performance problems. This view offers a high-level summary of key metrics such as max image size, oldest image age, total vulnerability count, and total running containers, providing a brief overview of the container image ecosystem at a glance. It also allows users to dive deeper into details about running containers associated with their images, track progress in remediating vulnerabilities, and use dashboards based on image-specific metrics for more customized monitoring. Additionally, it helps identify unused cloud registries and images taking up unnecessary disk space, enabling users to optimize their infrastructure footprint and costs.
May 06, 2024 874 words in the original blog post.
Datadog's Event Management is designed to help teams maintain service availability in complex cloud environments by aggregating and correlating events from disparate tools, reducing alert fatigue through AI-powered correlation and deduplication, enhancing alert context with service knowledge and ownership data, accelerating remediation with automated triage workflows, and integrating AIOps into monitoring workflows. By providing a unified view of incidents, teams can understand the full context of an issue, prioritize effectively, and respond more quickly and efficiently. This enables them to mitigate revenue and customer experience impacts associated with outages and improve overall incident response capabilities.
May 06, 2024 1,107 words in the original blog post.
Datadog Code Security has achieved an accuracy score of 100 percent in the OWASP Benchmark test, demonstrating its high dependability in detecting true threats and avoiding false positives. This is a significant milestone for an IAST solution, as traditional static application security testing tools often generate false positives and miss true risks. Datadog Code Security's production-ready approach is designed to provide performance without impacting application execution, making it suitable for real-world deployment. The tool offers over 20 additional detection rules beyond what is tested in the OWASP Benchmark, including server-side request forgeries, unvalidated redirects, hardcoded secrets, NoSQL injection attacks, and more. It provides highly specific information about detected vulnerabilities, empowering engineers to locate and fix them quickly. Datadog Code Security is a production-ready solution that allows teams to enjoy the advantages of IAST in a real-world setting, making it an attractive option for organizations seeking to improve their application security.
May 01, 2024 747 words in the original blog post.
The text discusses the development of a heatmap visualization tool for analyzing unaggregated data from multiple hosts in an infrastructure. The tool, powered by DDSketch, aggregates data during a flush interval and enables users to analyze statistical distributions across their entire infrastructure. The authors explore how they used DDSketch to build the Datadog heatmap visualization, including decisions on graphing distributions over time at an endless scale. They discuss the advantages of seeing unaggregated data, including the ability to visualize distinct modes without aggregation. The tool uses a color interpolation that aligns with the data's cumulative distribution curve and improves the overall contrast. Rendering is done by drawing rectangles on an HTML canvas, with performance improvements achieved by rendering per pixel instead of per data bin. The heatmap provides unique insights into system behavior, such as identifying changes in percentile values over time and unveiling seasonality and patterns in data that aggregations can hide. The tool's design and development process demonstrate the importance of a first-principles approach to building effective data visualizations at scale.
May 01, 2024 2,444 words in the original blog post.