July 2022 Summaries
24 posts from Datadog
Filter
Month:
Year:
Post Summaries
Back to Blog
Citrix Hypervisor, a type 1 hypervisor formerly known as Citrix XenServer, allows organizations to manage virtual infrastructures, including VMs, virtual desktops, and applications, with enhanced availability and flexibility through resource pools. Effective monitoring of these environments at all levels—from physical hosts to VMs and management toolstack—is crucial to maintaining performance and avoiding resource bottlenecks. Datadog's integration with Citrix Hypervisor facilitates this by collecting key metrics and logs that provide insights into resource utilization and potential bottlenecks. The integration allows users to monitor physical host resources, VM performance, session counts in resource pools, and XAPI memory usage, which are essential for optimizing virtual desktop infrastructures and ensuring efficient workload management. Users can track these metrics within Datadog's platform, which supports over 850 other services and technologies, offering a comprehensive monitoring solution.
Jul 29, 2022
960 words in the original blog post.
In this Datadog Spotlight Series, Benjamin Fernandes, Engineering Director based in Paris, shares his career growth journey at the company. He joined as an intern in 2013 and has since grown from an entry-level engineer to a director managing engineers across Europe for APM Distributed Tracing and Error Tracking. Datadog interns contribute significantly to the business, working on goals and projects similar to full-time employees. Fernandes emphasizes that Datadog's core values of being open, pragmatic, and honest have remained consistent as the company has grown. He encourages candidates considering joining Datadog by highlighting the opportunities for growth and exposure to a wide range of technologies in the global software engineering ecosystem.
Jul 28, 2022
757 words in the original blog post.
Performetriks, a service provider specializing in assessing and improving application performance and security, has released its Composer tool for Datadog on the Datadog Marketplace. The Composer tool allows users to store, track, and manage monitoring settings as code, making it easier to automate backups or deployment of configuration settings across multiple environments. It also streamlines coordination among administrators by enabling them to easily detect changes in configurations and create reusable, consistent monitoring setups. Additionally, the Composer tool allows for easy repair and restoration of Datadog configurations by maintaining backups of former settings. Performetriks' membership in the Datadog Partner Network enables it to promote its branded tools on the Marketplace, while interested developers can sign up as Technology Partners to create integrations or applications for the platform.
Jul 28, 2022
518 words in the original blog post.
Performetriks Composer is a tool that allows teams to streamline their application performance and security assessments for enterprise clients. It offers frameworks for automation, benchmarking, and security testing, as well as tools that evaluate and improve application performance. The Performetriks Composer for Datadog tool is now available in the Datadog Marketplace, enabling users to document and store Datadog configuration settings as code, making it easier to track and manage them. This allows admins to coordinate among multiple environments more efficiently, deploy uniform monitoring configurations, and easily repair and restore configurations in case of errors or changes. The tool promotes a monitoring-as-code solution using Datadog, providing benefits such as automating backups and restores, promoting branded monitoring tools, and enabling easy deployment of configuration settings.
Jul 28, 2022
529 words in the original blog post.
Sara Verdi here, and I'm excited to share with you the story of Benjamin Fernandes, an Engineering Director at Datadog who has grown his career from a gap year engineering internship to leading engineers across Europe for both APM Distributed Tracing and Error Tracking. Through his journey, Benjamin has learned to build large-scale distributed systems and work with customers and product managers to create useful features for the Datadog platform. He credits his growth to the company's core values of being open, pragmatic, and honest, as well as its supportive culture that allows employees to drive projects and learn from each other. As a testament to this, Benjamin shares stories of how he was able to build valuable systems as an intern, including a rate limiting system that is still running today, and how his management team supported him in his transition from individual contributor to manager. With Datadog's unique blend of global software engineering ecosystem exposure and high-stakes problem-solving, Benjamin advises candidates to consider joining the company for opportunities to grow as an engineer and learn from the best.
Jul 28, 2022
767 words in the original blog post.
Cilium is a Container Network Interface (CNI) that enhances security and load-balancing in Kubernetes environments by extending existing network capabilities. It allows teams to build advanced identity and application-aware network policies, replacing traditional firewalls and enabling cross-cluster communication. Cilium leverages Extended Berkeley Packet Filter (eBPF) to apply network and security logic in the Linux kernel without modifying application code or container configurations. Monitoring Cilium ensures that Kubernetes applications are processing requests as expected, making it a critical part of securing overall environments and supporting distributed applications.
Jul 25, 2022
2,966 words in the original blog post.
Hubble is a tool that enhances the monitoring of network traffic in Cilium-managed Kubernetes environments by providing both command-line and user interface capabilities. It collects and aggregates network data from all pods, offering insights into request throughput, status, and errors, and integrates with OpenTelemetry for exporting logs and traces. Hubble operates through servers and the Hubble Relay, which together monitor multiple levels of network traffic and make data accessible via APIs. While the Hubble CLI offers detailed network-level visibility, allowing users to identify issues like dropped requests and policy misconfigurations, the Hubble UI provides a high-level service map for understanding overall cluster interactions. Both interfaces use the same data points, enabling comprehensive troubleshooting and policy evaluation in Kubernetes networks.
Jul 25, 2022
1,018 words in the original blog post.
In this text, the author discusses how to use Datadog to integrate with Cilium's observability platform and gain end-to-end visibility into your network and Kubernetes environment. The integration enables visualization of Cilium metrics, analysis of logs for better insight into performance anomalies, monitoring of pod states using Live Container view, and observation of network traffic with Datadog's network performance and DNS monitoring tools. The author provides step-by-step instructions on enabling the integration via Kubernetes manifests and highlights key features that can help troubleshoot issues in your Cilium environment.
Jul 25, 2022
1,500 words in the original blog post.
Cilium is a Container Network Interface (CNI) provider that secures and load-balances network traffic in Kubernetes environments. It allows teams to build advanced identity and application-aware network policies, replacing traditional firewalls with enhanced security capabilities. Cilium uses Extended Berkeley Packet Filter (eBPF) to apply network and security logic in the Linux kernel without modifying application code or container configurations. This enables efficient resource utilization and reduces CPU overhead on worker nodes. Monitoring Cilium ensures that Kubernetes applications are processing requests as expected, making it a critical part of securing overall environments and supporting distributed applications. Key metrics include ipam_available, endpoint_state, policy_l7_total, cilium_api_limiter_processed_requests_total, and unreachable_nodes, which provide insights into IP address allocation, endpoint health, network policies, API processing, and node connectivity. By monitoring these metrics, teams can troubleshoot issues, optimize resource utilization, and ensure the overall security and performance of their Kubernetes environments.
Jul 25, 2022
3,082 words in the original blog post.
Datadog integrates with Cilium's observability platform to provide end-to-end visibility into network and Kubernetes environment metrics, enabling users to view service dependencies and traffic flows. The integration allows users to visualize Cilium metrics in Datadog's dashboard, analyze Cilium logs for better insight into performance anomalies, monitor the state of pods with Datadog's Live Container view, observe network traffic with Cloud Network Monitoring and DNS monitoring, and enable custom alerts for unusual log activity. By correlating data from Cilium with Kubernetes resources, users can identify potential issues before they become more serious, troubleshoot problems in their network, and gain a better understanding of their cluster performance.
Jul 25, 2022
1,442 words in the original blog post.
CockroachDB is a distributed SQL database developed by Cockroach Labs that ensures ACID semantics and easy horizontal scaling. It offers both self-hosted and fully-managed cloud-hosted versions, with the latter being CockroachDB Dedicated. Datadog integrates with CockroachDB to help monitor its performance and troubleshoot issues such as sudden drops in throughput or storage capacity. The integration collects and visualizes hundreds of metrics, allowing users to track overall database cluster health, workload, resource utilization, and more. By monitoring CockroachDB alongside other technologies like AWS and GCP, Datadog provides a comprehensive view of the entire tech stack for efficient troubleshooting and performance optimization.
Jul 22, 2022
672 words in the original blog post.
Powerpacks, a feature introduced by Datadog, are designed to enhance dashboard management and scalability within organizations by offering templated groups of widgets that can be saved and reused, thus promoting standardization and consistency across monitoring practices. These Powerpacks help organizations break down observability silos by enabling experts to capture and distribute their knowledge across different technologies and infrastructure components. By turning existing dashboard widgets into Powerpacks, teams can standardize monitoring for specific technologies or focuses, such as security, and ensure that key metrics are consistently tracked. Powerpacks can be customized with configuration variables to scope data to relevant contexts, and they allow for the quick and flexible creation of dashboards with both custom and out-of-the-box options available for Datadog products. This feature is aimed at improving the interpretability and consistency of dashboards, reducing mean time to recovery (MTTR) during incidents, and facilitating the scaling of monitoring knowledge as organizations grow.
Jul 20, 2022
913 words in the original blog post.
AWS Transit Gateway now supports VPC Flow Logs, allowing customers to gain deep visibility into network traffic across their Transit Gateways. This feature enables the capture of traffic through any or all attachments of a Transit Gateway and provides key information for troubleshooting, capacity planning, and security enhancements. Datadog has also released an integration that makes it easy to ingest and analyze these VPC Flow Logs for various use cases. By integrating with Datadog, users can more easily identify network issues, perform network capacity planning, and improve security by detecting suspicious activity or communication patterns.
Jul 14, 2022
1,073 words in the original blog post.
AWS Transit Gateway is a service that simplifies AWS networking architecture by eliminating the need to manage complex peering relationships and massive route tables. It improves security by ensuring that traffic between VPCs and Transit Gateways stays encrypted and avoids traveling over the public Internet. The newly announced support for VPC Flow Logs for Transit Gateway enables customers to capture deep, end-to-end visibility into all network traffic going through their Transit Gateways. This allows them to troubleshoot network issues, perform network capacity planning, and improve security by analyzing flow log records and identifying suspicious activity or communication patterns. Datadog now provides an integration that makes it easy to ingest and analyze VPC Flow Logs for Transit Gateway, enabling customers to leverage the benefits of this feature in their AWS environments.
Jul 14, 2022
1,128 words in the original blog post.
Google Cloud has introduced Arm-based Tau T2A virtual machines (VMs) for running workloads in Google Kubernetes Engine (GKE). Datadog provides complete visibility into GKE environments, including nodes running on Arm-based VMs. This is crucial for migrating workloads to Tau T2A machines as it enables users to compare costs and performance between Arm- and x86-based architectures. Datadog's built-in integration dashboard allows monitoring of key data across nodes, such as CPU utilization, to determine which ones could benefit from leveraging the new T2A machines. With Datadog APM, users can collect traces from nodes running on different architectures to compare performance data like request latency, error rate, and throughput. This helps ensure that newly provisioned nodes are appropriately configured to handle traffic as expected.
Jul 13, 2022
392 words in the original blog post.
Google Cloud has announced its Arm-based Tau T2A virtual machines (VMs) for running containerized workloads in Google Kubernetes Engine (GKE), providing a cost-effective and energy-efficient alternative to traditional architectures. This development is part of the growing trend of Arm processors in cloud computing ecosystems, with more organizations choosing to leverage their benefits. Datadog provides complete visibility into GKE environments, including Arm-based VMs, enabling users to easily compare costs and performance between different architectures. With Datadog's integration, users can automatically visualize data across their GKE clusters, analyze performance for Arm nodes, and start monitoring their Arm-powered workloads today.
Jul 13, 2022
405 words in the original blog post.
Service level objectives (SLOs) are crucial for maintaining service reliability and ensuring consistent value delivery to users. Adopting SLOs is an SRE best practice that helps teams balance priorities between remediating reliability issues and developing features. Two types of SLO alerts, error budget alerts and burn rate alerts, can provide ongoing visibility into service performance relative to objectives. Error budget alerts track the consumption of a service's error budget and notify teams when it passes a threshold, while burn rate alerts detect if the service is consuming its error budget more quickly than expected. These alerts help teams make informed decisions about priorities and address potential issues before they escalate.
Jul 05, 2022
2,007 words in the original blog post.
Service level objectives (SLOs) are used to measure a service's reliability and ensure it meets its performance goals. Adopting SLOs as an SRE best practice helps teams ensure their services perform well and consistently deliver value to users. To gain the greatest benefit from SLOs, teams need ongoing visibility into how well their services are performing relative to their objectives. Two types of alerts can be created: error budget alerts track consumption against a service's error budget, while burn rate alerts notify teams if they're consuming their error budget more quickly than expected. Burn rate alerts provide a proactive approach to SLO monitoring and can detect subtle changes that may impact user experience. Teams can choose from two approaches to decide on a burn rate alert threshold: estimating the time required to recover or setting a percentage of error budget consumption. By creating these alerts, teams can stay informed about issues that could deplete their error budget and ensure their services remain reliable.
Jul 05, 2022
2,028 words in the original blog post.
The Domain Name System (DNS) is crucial to internet functionality as it maps domain names to IP addresses. DNS-level events provide valuable information about network traffic that can be used to identify malicious activity, such as cryptojacking attempts and data exfiltration. Datadog's eBPF-powered Cloud Workload Security (CWS) now analyzes DNS activity in addition to file and process activity to detect security threats in real time. This enhances threat detection by providing visibility into DNS lookups, allowing users to spot attacks at the network level. The latest rules include out-of-the-box workload threat detection rules that flag suspicious activity like unexpected password changes, web shell creations, and nmap executions. Datadog CWS also includes rules for detecting "command and control" attacks and provides contextual information to help determine whether suspicious behavior is malicious.
Jul 01, 2022
698 words in the original blog post.
Datadog Summit, an in-person event celebrating community and knowledge sharing, will take place on August 16 in Sydney. Attendees can learn from engineers and developers who have transformed their organizations through observability cultures. Practical advice from community members using data and insights from Datadog to improve system performance, security, and reliability will be shared. The event will feature hands-on workshops covering infrastructure monitoring, distributed tracing, log management, and more. Registration is free but limited; RSVP now to reserve a seat at the Four Seasons Sydney.
Jul 01, 2022
483 words in the original blog post.
Azure Functions is an event-driven serverless compute service built on Azure App Service that allows developers to deploy code without managing infrastructure. To ensure optimal performance of functions, Datadog has released an extension for Azure App Service which collects traces and correlates telemetry from resources running in Azure App Service, including support for Azure Functions. This extension provides function-level visibility into application performance, allowing users to identify bottlenecks and troubleshoot issues. With Datadog APM, users can get an end-to-end view of request traces across their infrastructure, automatically tagging traces by function name. Additionally, the extension enables correlationating Azure Functions trace data with metrics, logs, and other traces from across Azure-hosted resources for deeper visibility into the health and performance of App Service plans.
Jul 01, 2022
482 words in the original blog post.
Azure Functions is an on-demand serverless compute offering built on top of Azure App Service that enables deployment of event-driven code without provisioning and managing infrastructure. To ensure applications respond quickly when invoked, it's essential to monitor performance and resource bottlenecks. The Datadog extension for Azure App Service collects traces and correlates telemetry from resources running in Azure App Service, providing function-level visibility into serverless Azure infrastructure. This allows users to identify performance issues, optimize their Azure-hosted applications, and troubleshoot problems efficiently by breaking down request traces into spans and automatically tagging them with function names. The Datadog extension also enables correlation with metrics and logs from across the Azure-hosted resources, helping users determine whether performance issues are related to underlying capacity problems or other factors. By integrating the Datadog App Service extension with Azure Functions, users can gain deeper visibility into their serverless infrastructure and take steps to reconfigure their functions and App Service plans for better performance.
Jul 01, 2022
493 words in the original blog post.
Datadog's Cloud Workload Security (CWS) now analyzes DNS activity in addition to file and process activity to detect security threats in real time. This new feature provides visibility into DNS lookups, enabling detection of malicious activity such as cryptojacking attempts and data exfiltration. CWS includes out-of-the-box workload threat detection rules that flag suspicious activity, including unexpected password changes, web shell creations, and nmap executions. The platform also correlates related security signals to provide contextual information, helping users determine whether suspicious behavior is malicious. With the addition of DNS-based threat detection, Datadog's CWS provides another layer of protection for environments, detecting threats at the network level and providing a more comprehensive view of security posture.
Jul 01, 2022
711 words in the original blog post.
The Datadog Summit is a celebration of community, where engineers and SREs can share knowledge and learn from others in the community. The event will feature talks from users who have transformed their organizations by building cultures of observability, as well as hands-on workshops covering topics such as infrastructure monitoring, distributed tracing, log management, and more. These workshops will provide real-world insights on how to understand system performance and quickly find and resolve issues. The Datadog team will also be present to show off new product features and answer questions. The summit will take place in Sydney on August 16 at the Four Seasons Sydney, with limited space available, so RSVP is recommended.
Jul 01, 2022
493 words in the original blog post.