October 2026 Summaries
13 posts from Datadog
Filter
Month:
Year:
Post Summaries
Back to Blog
No summary generated yet.
Oct 09, 2026
956 words in the original blog post.
No summary generated yet.
Oct 08, 2026
1,273 words in the original blog post.
No summary generated yet.
Oct 08, 2026
1,099 words in the original blog post.
Datadog migrated its host vulnerability scanning from a remote SSH-based scanner to the Datadog Agent after fleet growth made it difficult to consistently assess all hosts before scanner credentials expired. The Agent-based approach, integrated with existing infrastructure and observability workflows, became the system of record for vulnerability findings and maintained scan freshness above 99 percent for in-scope hosts within 24 hours. During a six-month parallel evaluation, Datadog compared coverage, package identification, vulnerability detection, reporting, audit evidence, and remediation workflows, while investigating differences such as software classification, handling of unloaded kernels, and distinct methods of grouping CVEs. The migration enabled teams to prioritize vulnerabilities using production context, address shared issues through image and infrastructure release processes, and support compliance obligations including SOC 2 Type II, ISO 27001, PCI DSS, and FedRAMP High. Following cutover, Datadog retired its dedicated scanners and remote credentials, emphasizing that comparable migrations should define existing controls, establish replacement requirements, and validate meaningful differences through parallel operation.
Oct 07, 2026
1,588 words in the original blog post.
Datadog CI/CD Optimization is presented as a platform for helping DevOps, platform, and developer-experience teams manage the increased CI workload associated with AI-assisted software development by improving speed, reliability, and cost efficiency. It combines pipeline, test, commit, log, runner, and infrastructure data to identify whether delays stem from queue times, slow jobs, regressions, or resource constraints, while monitoring can alert teams to pipeline failures and performance degradation. Its flaky-test capabilities prioritize costly intermittent failures, automatically retry failed tests, and use detection and pull-request gates to prevent known or newly introduced flakes from blocking delivery. Test Impact Analysis skips tests unrelated to a code change, and Test Parallelization distributes necessary tests based on expected duration to reduce completion time and runner use. The platform centralizes pipeline and test investigation across CI providers including GitHub Actions, GitLab CI/CD, Jenkins, Azure DevOps, CircleCI, and Buildkite, enabling teams to correlate CI issues with infrastructure data and process more development volume without proportional growth in compute costs.
Oct 06, 2026
1,509 words in the original blog post.
Datadog Observability Pipelines Exabeam Packs are preconfigured, source-specific filters designed to reduce the volume and cost of security logs sent to Exabeam SIEM while preserving the raw payloads and detection-relevant fields required by its parsers, UEBA models, and threat rules. The Packs apply dropping, deduplication, and sampling to routine activity from network, endpoint, web, identity, and collaboration sources, retaining higher-value events such as denied connections, VPN activity, security alerts, suspicious processes, Sysmon and PowerShell events, and blocked web traffic. Supported sources include Cisco ASA, Fortinet FortiGate, Palo Alto, CrowdStrike Falcon, SentinelOne, Windows, and Zscaler, with additional destination-agnostic coverage for Active Directory, Okta, Microsoft 365, and Abnormal.ai. Organizations can validate filtering through Live Capture, retain complete copies of logs in cloud archives such as Amazon S3, generate metrics from filtered events, and replay selected archived data into Exabeam for audits or investigations. The approach aims to provide a lower-noise SIEM stream without losing long-term visibility, and it can also route data to other destinations, including Datadog Cloud SIEM.
Oct 05, 2026
1,642 words in the original blog post.
A hypothetical ransomware attack illustrates how the first 72 hours require rapid technical containment, business coordination, and communication planning, beginning with activating incident response procedures, engaging forensic experts, and using offline crisis documentation. Organizations should preauthorize personnel to isolate affected systems, investigate compromised identities and malware spread, verify that backups are clean, and maintain alternative communication channels for employees, customers, media, and leadership when normal tools fail. Although protected, tested backups enable restoration without relying on attackers, full recovery can take months because teams must rebuild dependencies, validate data, prioritize systems, clear operational backlogs, and restore customer trust. Legal, regulatory, law-enforcement, and board-reporting responsibilities should be assigned in advance, particularly because data theft can create notification obligations and ransom payments do not guarantee decryption or deletion of stolen information. The piece also argues that AI-assisted security investigation, combining observability and security data, can accelerate detection and triage of early attacker activity, while tools such as Datadog Cloud SIEM can correlate signals, prioritize risky entities, support automated investigation, and enable proactive threat hunting.
Oct 02, 2026
2,120 words in the original blog post.
Databricks provides several complementary sources of telemetry for monitoring data engineering, analytics, and Model Serving workloads, with system tables serving as the primary resource for analyzing performance, costs, usage, audit activity, and lineage across Unity Catalog-enabled workspaces. Because system-table data is not real time, live monitoring relies on options such as job notifications, APIs, host-based agents, cloud-provider telemetry, and Model Serving metrics endpoints. Compute monitoring includes system tables for clusters, nodes, instances, and SQL warehouses, while classic clusters can also send logs and metrics externally and be investigated through the Spark UI for execution timelines, DAGs, skew, spill, shuffle, streaming behavior, and driver or executor logs. Job and pipeline performance can be assessed through Lakeflow system tables, Databricks job and pipeline interfaces, event logs, and Spark Streaming Query Listener, with notifications available for immediate status and duration alerts. Query performance is available through query-history tables, APIs, UI profiling, and serverless query insights, while data quality and lineage capabilities include Unity Catalog profiles, anomaly detection, SQL alerts, audit logs, lineage tables, Catalog Explorer graphs, and optional OpenLineage integration. For Model Serving, Databricks offers endpoint health metrics, logs, inference tables, model-quality profiles, LLM evaluation tools, and AI Gateway usage data to analyze request volume, errors, latency, compute consumption, model behavior, and costs.
Oct 02, 2026
1,434 words in the original blog post.
Databricks is a lakehouse platform that combines open data formats, scalable object storage, ACID transactions, governance, analytics, AI, and data-management capabilities in shared workspaces divided between control and compute planes. Its workloads include scheduled jobs, Spark Declarative Pipelines, SQL warehouses, streaming applications, and Model Serving endpoints, running on either customer-managed classic compute or Databricks-managed serverless compute. Effective monitoring relies on control-plane system tables, Spark execution signals, query telemetry, infrastructure metrics where host access is available, and endpoint metrics to identify failures, latency, resource saturation, data-processing backlogs, and autoscaling limitations. Important indicators include job outcomes and durations, pipeline quality errors, task failures, shuffle spills, garbage-collection time, warehouse queue waits and query latency, streaming input-versus-processing rates, CPU and memory use, inference error rates and tail latency, and GPU memory pressure. Because Databricks billing is based on elastic per-second DBU consumption, tracking usage and estimated cost is also necessary to detect anomalies caused by misconfigured infrastructure, retries, or inefficient workloads.
Oct 02, 2026
2,208 words in the original blog post.
Datadog’s Databricks integration combines Agent-based Spark and infrastructure telemetry, API polling for job execution data, and system-table queries for cost and lineage information to provide visibility across analytics, data engineering, and AI/ML workloads. Its Jobs Monitoring features track job and cluster health, performance, logs, traces, resource utilization, failures, delays, and costs, while offering anomaly detection, alerting, troubleshooting tools, and rightsizing recommendations that can be linked to cloud infrastructure telemetry and Jira workflows. Data Observability Quality Monitoring extends coverage to Delta and Unity Catalog table freshness, volume, column metrics, custom rules, and lineage, while Data Streams Monitoring maps streaming pipelines and helps correlate upstream issues such as Kafka lag with Spark job failures. The integration also monitors Model Serving endpoints through latency, throughput, errors, CPU, memory, and GPU metrics, with dashboards and alert templates for detecting scaling or reliability problems. Datadog Cloud Cost Management contextualizes Databricks DBU spending with discounted cloud costs, tag-based attribution, anomaly detection, forecasts, and custom metrics for analyzing costs of training, ETL, and inference workloads. Reference Tables further enrich logs, metrics, and events with business and operational metadata to support ownership, cost allocation, targeted alerts, and configuration analysis.
Oct 02, 2026
1,590 words in the original blog post.
Datadog created Distributed DataFusion, an open-source Rust framework that extends Apache DataFusion from single-machine execution to low-latency distributed query processing across multiple machines. The project supports Datadog’s move toward a unified, composable data stack based on Apache Arrow, Substrait, and DataFusion, reducing reliance on separate specialized engines for metrics, logs, traces, and other observability data. Unlike existing distributed systems such as Spark, Trino, ClickHouse, and Ballista, it was designed for extensibility, Arrow and Substrait integration, zero-copy streaming, and efficient interactive workloads. It distributes DataFusion physical plans into stages and tasks, using mechanisms such as NetworkShuffle to repartition data between workers for operations like grouped aggregations and NetworkCoalesce to collect distributed results. The framework can retain lightweight queries on one machine to avoid network, serialization, and coordination costs, while enabling heavy queries to run across workers; Datadog reports some production queries complete up to 10 times faster at the same overall cost. Distributed DataFusion is available under the Apache 2.0 license and is being further developed with adaptive query execution and potential GPU acceleration, while allowing users to supply their own data sources, networking, execution nodes, and planning logic.
Oct 01, 2026
2,724 words in the original blog post.
Datadog has expanded its open-source Supply Chain Firewall through Code Security to provide organization-wide protection against malicious software packages before they are installed. The firewall intercepts supported npm, pip, and poetry commands, evaluates proposed dependencies against malicious-package intelligence, vulnerability advisories, and recency checks, and can block or warn developers about risky packages. New centralized management enables organizations to apply shared allowlists and blocklists across developer environments and collect firewall events in a unified reporting feed. Code Security also retrospectively reevaluates recorded installations as threat intelligence changes, alerting teams when packages previously installed on workstations or CI systems are later identified as malicious. A new GitHub Action extends the same inspection, policy enforcement, and reporting to CI runners, addressing package-installation risks across local development, automated pipelines, and AI-assisted development workflows.
Oct 01, 2026
766 words in the original blog post.
Successful data pipelines can still deliver stale, incomplete, duplicated, or malformed warehouse data, making pipeline execution monitoring insufficient on its own. The discussion recommends monitoring data at rest through table-level freshness and row-count checks, column-level measures such as nullness, uniqueness, cardinality, and distribution, schema-change detection, and custom SQL rules for business-specific conditions. It contrasts anomaly detection, which learns historical trends and seasonality but needs training data, with fixed thresholds for stable requirements or explicit service-level rules. To manage coverage and alert noise, teams should prioritize critical tables, group monitors appropriately, account for query costs, and route alerts to both data owners and pipeline responders. When issues arise, lineage and telemetry can help trace downstream impact and distinguish source, batch-processing, streaming, or hybrid-architecture causes by comparing input and output volumes, throughput, lag, logs, query history, and infrastructure signals. Datadog Data Observability and related monitoring tools are presented as a way to connect warehouse data-quality alerts with pipeline and stream context for faster investigation.
Oct 01, 2026
2,273 words in the original blog post.