November 2020 Summaries
27 posts from Datadog
Filter
Month:
Year:
Post Summaries
Back to Blog
The 2020 North America edition of CNCF's KubeCon + CloudNativeCon marked its fifth anniversary with 25,000 registrants. As Kubernetes has entered the early majority phase, discussions at the conference focused less on development and more on end-user migration stories, security, developer experience, and cloud native ecosystem projects. Developer experience was a major topic, with emphasis on improving configuration management and developer workflows. Security and governance were also important topics, as Kubernetes enters the enterprise sector. Multi-cluster management was another focus area, with solutions being developed to manage multiple clusters effectively and securely.
Nov 30, 2020
1,153 words in the original blog post.
Datadog has been selected as Customers' Choice for Application Performance Monitoring in Gartner Peer Insights, a peer review platform that recognizes vendors with high customer ratings. The company's application performance monitoring (APM) solution is praised by customers for its ability to monitor and optimize end-to-end application performance in a single pane of glass. Datadog APM has been described as the "Swiss army knife of software monitoring" and an "excellent partner who is open to business needs and feature developments." The solution provides unparalleled visibility into modern applications through various offerings, including App Analytics, Watchdog, and the Service Map. Customers appreciate Datadog's ability to respond to feedback and deliver a stronger user experience.
Nov 30, 2020
300 words in the original blog post.
KubeCon + CloudNativeCon is the most important event for Kubernetes adopters and technologists, with a focus on end-user migration stories, security, developer experience, and the cloud native ecosystem. The conference has matured, with enterprises embracing Kubernetes and cloud native technologies, and industry leaders discussing topics such as developer experience, configuration management, GitOps, developer workflows, security, governance, multi-cluster management, and emerging trends like machine learning, edge computing, and network functions virtualization. Industry projects like SPIFFE, Spire, Checkov, and Open Policy Agent are gaining traction, offering solutions for shift-left security, DevSecOps, and policy as code. The conference has seen the emergence of new solutions for multi-cluster workload management, including cluster scheduling APIs, and Datadog continues to participate in the event to ensure its monitoring tool remains the best fit for Kubernetes clusters and workloads.
Nov 30, 2020
1,167 words in the original blog post.
AWS re:Invent, an annual conference held in Las Vegas since 2012, is going virtual this year due to the pandemic. The event will be spread across three weeks instead of one and sessions will be rebroadcast to serve multiple time zones. Datadog, a long-time AWS partner, has prepared various interactive activities for attendees including demos in four languages, 1:1 sessions with technical solutions team members, daily raffle for an Xbox Series X, swag passport program and free trial sign up sweepstakes for a chance to win an iPhone 12 or Google Pixel 5. A virtual scavenger hunt through the Datadog platform will also be hosted on Tuesday, December 8. New for 2020 are AWS Hack Jams where attendees can compete in solving hands-on real-world challenges in live AWS environments. The event will feature several sessions including building the next generation of residential robots, paving the way toward automated driving with BMW Group, deep dive into AWS Lambda security: function isolation, testing resiliency using chaos engineering and Adrian Cockcroft’s architecture trends and topics for 2021. Datadog will also deliver a session with one of its customers, Dunelm, on how they drastically improved their deployment velocity by moving to serverless everything.
Nov 24, 2020
1,398 words in the original blog post.
Waldo Grunenwald, an AWS partner, looks forward to re:Invent every year, enjoying the opportunity to meet face-to-face and share the latest in monitoring and security. This year's event will have a virtual format, spanning three weeks instead of one week, with sessions rebroadcasted for multiple time zones. The event features engaging activities such as Datadog demos, 1:1 sessions with technical solutions team members, a daily raffle for an Xbox Series X, and a swag passport program. Notable sessions include "Building the next generation of residential robots," "Paving the way toward automated driving with BMW Group," "Deep dive into AWS Lambda security: Function isolation," and "Testing resiliency using chaos engineering." Datadog's take on these topics includes discussions on automata, autonomous driving data lakes, serverless first architectures, continuously tested resilience, and sustainability concerns. The event also features a session by Adrian Cockcroft highlighting emerging architecture trends and topics for 2021.
Nov 24, 2020
1,422 words in the original blog post.
Datadog's Cloud Workload Security provides real-time threat detection for production workloads in cloud environments. It monitors file, process, and kernel activity across the environment to detect threats at the infrastructure and workload levels. With Datadog Cloud Workload Security, developers can focus on threats holistically without sacrificing visibility or ease of management. The unified Datadog Agent is used to monitor the environment, and the platform combines real-time threat detection with metrics, logs, traces, and other telemetry from over 850 technologies. This allows teams to see the full context surrounding a potential attack and quickly investigate and respond to active threats in their cloud environment. The platform also provides Security Signals that contain critical context necessary for investigation, including key process metadata and MITRE ATT&CK tactics and techniques. Datadog's Cloud Workload Security view provides a full-picture perspective on the security posture of workloads, allowing teams to visually track Security Signals across their life cycle and respond faster with cross-stack correlation.
Nov 23, 2020
952 words in the original blog post.
Datadog Workload Protection, formerly known as Cloud Workload Security, offers real-time threat detection for production workloads by monitoring file, process, and kernel activity across environments such as AWS EC2 instances and Kubernetes clusters. This tool allows developers and security teams to maintain full-stack visibility without sacrificing ease of management, enabling them to address threats holistically. Integrated into the broader Datadog platform, Workload Protection provides comprehensive security by combining real-time threat analysis with metrics, logs, and traces from over 900 technologies. It uses out-of-the-box Detection Rules to identify malicious activities, such as web shell attacks, and generates Security Signals that include critical context for investigation, which are mapped to MITRE ATT&CK tactics. Furthermore, Datadog allows users to correlate runtime events with application logs, facilitating swift investigation and response. Existing Datadog users can immediately integrate Workload Protection, while new users have the option of a 14-day free trial.
Nov 23, 2020
971 words in the original blog post.
To ensure optimal performance in a vSphere environment, it is crucial to monitor key metrics that provide insight into resource usage and overall health. These metrics include CPU utilization, memory usage, disk I/O, network throughput, and tasks and events. By setting up alerts for these metrics, you can proactively identify potential issues before they impact your virtual machines (VMs) or applications running on them.
In this guide, we will discuss the importance of monitoring key vSphere metrics and provide examples of how to interpret their values. We will also cover some common causes of performance degradation in a vSphere environment and suggest ways to address them.
CPU Metrics:
1. CPU usage (%): This metric measures the percentage of CPU resources that are currently being utilized by all VMs running on an ESXi host. If this value consistently exceeds 80%, it may indicate that your VM configuration is not optimized, or that there is a resource contention issue between multiple VMs.
2. CPU ready (%): This metric represents the percentage of time that a VM has been waiting for access to the underlying physical CPU resources on an ESXi host. High values (e.g., > 5%) can indicate that your VM configuration is not optimized, or that there is a resource contention issue between multiple VMs.
3. CPU core usage (%): This metric measures the percentage of CPU cores that are currently being utilized by all VMs running on an ESXi host. If this value consistently exceeds 80%, it may indicate that your VM configuration is not optimized, or that there is a resource contention issue between multiple VMs.
Memory Metrics:
1. Memory usage (%): This metric measures the percentage of physical memory resources that are currently being utilized by all VMs running on an ESXi host. If this value consistently exceeds 80%, it may indicate that your VM configuration is not optimized, or that there is a resource contention issue between multiple VMs.
2. Balloon driver (vmmemctl) capacity: Each VM in vSphere can have a balloon driver (named vmmemctl) installed on it. If an ESXi host runs low on physical memory that it needs to allocate, it can reclaim memory from the guest physical memory of virtual machines by sending requests to the balloon drivers to "inflate" by gathering unused memory from the VM. The ESXi host can take that memory from the "inflated" balloon driver and deallocate the appropriate mapped host physical memory, which it can then allocate to other VMs. This technique is known as memory ballooning.
3. Memory swapped in/out: When an ESXi host provisions a virtual machine, it allocates physical disk storage files known as swap files. Swap file size is determined by the VM's configured size, less any reserved memory. For instance, if a VM is configured with 3 GB of memory and has a 1 GB reservation, it will have a 2 GB swap file. By default, a VM's swap files are collocated with its virtual disk, on shared storage.
4. Active memory versus consumed memory: In order for a VMKernel to accurately discern how much memory is actively in use by VMs, it would need to monitor every memory page that has been read from or written to. This process, however, would require too much overhead. Instead, the VMKernel uses algorithmic learning to generate an estimate of each VM's active memory usage.
5. Memory usage: At the VM level, the mem.usage metric measures what percentage of its configured memory a VM is actively using. Ideally, a VM should not always be using all of its configured memory. If it is consistently using a large portion of its configured memory, the VM will be less resilient to any spikes in memory usage if its ESXi host cannot allocate additional memory.
Disk Metrics:
1. Disk commands aborted: In vSphere, a single storage device cluster may hold datastores that serve many virtual machines. If there is a surge of commands from virtual machines to the storage hardware where datastores are located, that storage may become overloaded and unresponsive.
2. Disk bus resets: If a storage device is overwhelmed with too many read and write commands from an ESXi host, or if it encounters a hardware issue and fails to abort commands, it will clear out all commands waiting in its queue. This is called a disk bus reset.
3. Datastore provisioned capacity and actual VM usage: Storage is a finite resource. The diskspace.provisioned.latest metric tracks how much storage space is available on the datastores that the ESXi host communicates with, while virtualDisk.actualUsage lets you monitor how much disk space the VMs running on that host are actively using.
4. Disk latency: Monitoring latency is key to ensuring that your VMs are communicating with their virtual disks efficiently and without delay. Total disk latency measures the time it takes, in milliseconds, for an ESXi host to process a request sent from a VM to a datastore.
5. Queue latency: Depending on their configuration, storage devices like LUNs have a limited number of commands they can queue at any one time. When the volume of virtual machine commands sent from an ESXi host exceeds what a storage device can queue itself, those commands will begin to queue in the VMKernel.
6. Disk throughput: To ensure that your datastores, ESXi hosts, and VMs are processing read and write commands without interruption, monitor their I/O throughput for visibility into their activity.
Network Metrics:
1. Network received and network transmitted: These metrics track the network throughput, in kilobytes per second, of the object you're observing whether it's a host or a VM.
Tasks and Events:
By default, vSphere records tasks and events that occur in the VMs, ESXi hosts, and the vCenter Server of your virtual environment. These can include user logins, VM power-downs, certification expirations, and host connects/disconnects. Monitoring the events in these log files can help you stay aware of overall activity within your vSphere clusters and also perform audits and investigate any issues that occur in your environment.
In conclusion, monitoring key metrics in a vSphere environment is essential for maintaining optimal performance and ensuring that resources are being utilized efficiently. By setting up alerts for these metrics and regularly reviewing their values, you can proactively identify potential issues before they impact your virtual machines or applications running on them.
Nov 19, 2020
6,113 words in the original blog post.
In this post, we discuss how to access key VMware vSphere metrics using built-in monitoring tools like the vSphere Client and esxtop command-line tool. We also cover configuring vSphere to use a syslog forwarder for long-term log storage and analysis. Additionally, we explore data collection intervals and levels in vSphere, which help control the volume of collected data. The vSphere Client allows administrators to visualize key metrics with performance charts, set alarms on metrics, view tasks and events, and export logs. Esxtop provides more granular monitoring capabilities for virtual machines and hosts. Finally, we discuss configuring log forwarding from ESXi hosts and the vCenter Server to an external syslog server.
Nov 19, 2020
2,621 words in the original blog post.
In this text, the author discusses how to use Datadog for complete end-to-end visibility into physical and virtual layers of vSphere environments. The integration collects metrics, traces, logs, and more into a single platform, allowing users to monitor the health and performance of their vSphere environment as well as applications and services running on VMs. The author provides step-by-step instructions for enabling Datadog's vSphere integration, configuring metric collection, visualizing key metrics with dashboards, collecting logs, and monitoring applications running in virtual environments. Additionally, the text highlights how Datadog's logging features can help manage large volumes of logs and how users can monitor their entire stack alongside vSphere using more than 650 integrations.
Nov 19, 2020
2,468 words in the original blog post.
VMware vSphere is a virtualization platform that allows users to provision and manage one or more virtual machines on individual physical servers using the underlying resources. It enables organizations to optimize costs, centrally manage their infrastructure, and set up fault-tolerant virtual environments. The platform uses Distributed Resource Scheduler (DRS) and vMotion to automatically distribute shared physical resources to VMs based on their needs and ensure zero downtime during maintenance or when a server is overburdened. Monitoring performance and capacity management are crucial in vSphere, as poor resource allocation can lead to degraded performance or even downtime. Key metrics include summary metrics for high-level insight into infrastructure size and health, CPU usage, memory usage, disk metrics, network metrics, and tasks and events. These metrics help administrators track the status of their virtual environment, identify potential issues, and make informed decisions to optimize resource allocation and ensure optimal performance.
Nov 19, 2020
6,320 words in the original blog post.
Datadog's vSphere integration collects metrics and events from your vCenter Server, allowing you to monitor the health and performance of your vSphere environment in a single platform. The integration is built into the Datadog Agent, which can be installed on a single VM connected to your vCenter Server to collect data from all ESXi hosts, virtual machines, clusters, datastores, and data centers. To enable the integration, you'll need to configure the Agent's vSphere integration, including setting up role-based permissions and configuring the collection level and resource filters. Once enabled, Datadog will automatically start collecting monitoring data and populating out-of-the-box dashboards with key metrics, including CPU usage, memory ballooning, and disk latency. The platform also includes features such as event stream widgets, tags, alerts, and log forwarding to provide a unified view of your vSphere environment. Additionally, Datadog's APM provides full visibility into the performance of applications running on your virtual machines by collecting distributed traces, logs from VM-hosted applications, and more than 850 integrations across various technologies and services. With this integration, you can quickly understand the health and performance of your infrastructure, scale your vSphere resources appropriately, and identify and troubleshoot issues in a single platform.
Nov 19, 2020
2,419 words in the original blog post.
AWS Network Firewall is a firewall designed for Amazon Virtual Private Cloud (VPC) with support for third-party intrusion detection systems like Snort and Suricata. Datadog has partnered with AWS to integrate its platform with the firewall, allowing users to monitor firewall traffic within their VPC network and detect potential threats. The integration helps users understand firewall performance, discover trends in firewall flow logs, and use the firewall to detect threats. With this integration, users can visualize AWS Network Firewall metrics alongside VPC metrics for a comprehensive view of network security.
Nov 17, 2020
678 words in the original blog post.
Oracle Cloud Infrastructure (OCI) is an IaaS and PaaS cloud service used by large companies for hosting, storage, networking, and more. Datadog now integrates with OCI Logging to provide a comprehensive view of users' cloud environments. The integration allows users to stream logs directly into Datadog, where they can be stored indefinitely, analyzed for troubleshooting, and monitored for security and compliance posturing. With the integration, users can filter and search for important events within their OCI environment, visualize log data in metrics dashboards, detect and alert on security vulnerabilities, and create custom dashboards to visualize key log data from their environment. Additionally, Datadog Cloud SIEM provides a centralized location for teams to detect and triage security threats across the entire infrastructure stack.
Nov 17, 2020
1,029 words in the original blog post.
AWS Network Firewall is a firewall solution for Amazon Virtual Private Cloud (VPC) that provides pluggable support for third-party intrusion detection systems. It helps customers monitor and guard against unwanted traffic to and from their VPCs, with fine-grained visibility into potential attacks. Datadog's integration with AWS Network Firewall enables users to understand performance, discover trends in firewall traffic, quickly detect threats, and visualize metrics alongside other network infrastructure running in the VPC. The integration also makes it easy to spot unexpected surges or drops in traffic and provides an anomaly monitor to alert teams automatically if issues arise.
Nov 17, 2020
692 words in the original blog post.
Oracle Cloud Infrastructure (OCI) users can now integrate their OCI logs with Datadog for enhanced monitoring and security. The integration allows OCI users to stream all types of logs, including audit logs, service logs, and custom logs, directly into Datadog, where they can be stored indefinitely, analyzed for troubleshooting, and monitored for security and compliance posturing. With the integration, OCI users can use key event metadata to filter and search for important events, visualize their log data in metrics dashboards, detect and alert on security vulnerabilities, and create custom dashboards to track key log data from their environment. The integration also provides a centralized location for detecting and triaging security threats with Datadog Cloud SIEM.
Nov 17, 2020
1,037 words in the original blog post.
Microsoft Azure, a rapidly expanding cloud platform, has enhanced its integration with Datadog, a monitoring service, to provide more efficient and seamless monitoring capabilities. Datadog has reduced latency for Azure Monitor metrics ingestion by over 40%, ensuring faster and more accurate updates on performance metrics, which is crucial for tracking changes in deployment failures and resource limits. The integration now includes a streamlined log collection process through an easily deployable template, allowing for quick setup and configuration of Azure logs into Datadog. Additionally, Datadog has introduced new out-of-the-box dashboards for Azure CosmosDB and Azure API Management Services, enabling users to monitor these services more effectively with instant insights into performance and usage metrics. These enhancements aim to provide comprehensive visibility and management tools for Azure users, with the updates available immediately and a 14-day free trial offered for new users.
Nov 17, 2020
629 words in the original blog post.
KubeCon North America 2020 is set to be a major event for adopters and technologists in the Kubernetes community, with numerous keynotes, announcements, and sessions lined up. This year's virtual conference presents challenges in terms of focus and engagement, but it also offers opportunities to learn from experts worldwide.
Datadog has shared its list of recommended sessions for this year's event, including topics such as static analysis of Kubernetes manifests, enhancing the Kubernetes scheduler for diverse workloads, writing a kubelet in Rust, and managing Kubernetes in regulated environments. The company will also be hosting several events during the week of KubeCon, including a session on Kubernetes monitoring, a workshop on autoscaling applications deployed to Kubernetes, and two sessions featuring its engineers discussing their experiences with Kubernetes at scale.
In addition to these recommendations, Datadog encourages attendees to explore the full schedule for more interesting sessions and activities. The company will be hosting a virtual booth during the event and offering opportunities to win prizes such as an Xbox Series X or an iPhone 12 Pro.
Nov 13, 2020
1,965 words in the original blog post.
Amazon CloudFront is a content delivery network (CDN) that minimizes latency by caching your content on AWS edge locations around the world. With CloudFront real-time logging, you can understand how efficiently CloudFront is distributing your content and responding to requests. You can collect CloudFront real-time logs in Datadog—in addition to CloudFront metrics—to get deep visibility into the health and performance of your CloudFront distribution.
In this post, we’ll show you how to configure CloudFront to send real-time logs to Datadog, and how to organize, analyze, and alert on your logs. From CloudFront to Datadog, CloudFront sends real-time logs to Amazon Kinesis, a managed streaming data service. You can configure Kinesis to forward your logs to a destination of your choice, such as Datadog.
To create your log configuration, provide a name for the configuration and specify its log sampling rate—the percentage of logs generated by CloudFront that you want to send to Kinesis. Next, select the fields to include in your logs. By default, your log configuration will include all of the available CloudWatch log fields, as shown in the screenshot below. You can easily configure Datadog to parse this format automatically.
If you only want to log a subset of the available fields—for example, to reduce the amount of data in your stream—you can modify the log configuration’s Fields list and then create a custom log pipeline to parse your modified log format. Next, designate a Kinesis Data Stream as the endpoint to which CloudFront will send your logs. If you already have a stream you want to use, enter its Amazon Resource Name (ARN) in the Endpoint field. Or to create a new stream, click the Kinesis link, then fill in the fields shown in the screenshot below.
Once Kinesis has created your data stream, copy its ARN as shown here, then navigate back to the Create real-time log configuration page and paste the ARN into the Endpoint field. To finish creating your real-time log configuration, specify the IAM role CloudFront will use to send your logs to Kinesis. Finally, select the CloudFront distribution and cache behaviors that will generate the logs, and click Create configuration.
Kinesis Data Firehose is a managed service that can route streaming data in near real time to AWS services, HTTP endpoints, and third-party services like Datadog. Kinesis Data Firehose makes it easy to stream AWS service logs into Datadog—including real-time logs from CloudFront.
To route your logs into Datadog, create a Kinesis Data Firehose delivery stream and choose the Kinesis Data Stream you created above as the source: Next, specify Datadog as your delivery stream’s destination, as shown in the screenshot below. Finally, enter your Datadog API key, select the appropriate HTTP endpoint URL, and provide the required additional configuration details.
Once Datadog is ingesting your CloudFront real-time logs, you can use the Log Explorer to view, search, and filter your logs to better understand the performance of your CloudFront distribution. In this section, we’ll show you examples of CloudFront log fields you can use to investigate the source of errors and latency. But first, we’ll explain how you can use tags to make it easy to explore your CloudFront logs.
Tags give you the ability to group and filter your logs on multiple dimensions and correlate them with data from other services you’re monitoring. Datadog automatically applies AWS tags like region and aws_account to your CloudFront logs, and you can add your own tags to associate your logs with related metrics from CloudFront and other AWS services like Amazon RDS or Elastic Load Balancing.
To apply a custom tag to your CloudFront logs, create a parameter on your Kinesis Data Firehose delivery stream. In the screenshot below, we’ve added a parameter with a key of distributionid so we can isolate log data by distribution. If you’ve created a Kinesis Data Firehose delivery stream for each of your CloudFront distributions, then you can use the distributionid tag to easily correlate your logs with metrics from the same distribution.
Now that you’re collecting and tagging your real-time CloudFront logs, you can analyze them in Datadog to monitor your distributions’ performance. You can create facets and measures based on the data in your logs to help you improve user experience by tuning parameters like cache size and time to live (TTL). For example, the x-edge-response-result-type field in your logs shows values such as Hit and Miss that describe how an edge location responded to each request.
You can also create log-based metrics to monitor long-term trends in your logs. Like other metrics, log-based metrics are retained for 15 months at full granularity. In the screenshot below, we’ve created a log-based metric to track latency experienced by users in different geographies. To generate the metric, we queried the time-to-first-byte log field—which tells how long CloudFront took between receiving the request and beginning to respond—and grouped it by a log facet we created on the edge-location field.
We can graph this metric to see the latency values logged by each of the edge locations used by our distributions, as shown in the screenshot below. Alerts Searching and filtering your logs is a powerful tool for investigation, but you can also use log monitors to proactively notify your team any time your CloudFront logs indicate a potential problem. The screenshot below shows an example of a log monitor that watches for an increase in the rate of 502 errors, which can occur when CloudFront is unable to communicate with your distribution’s origin server.
Nov 13, 2020
1,196 words in the original blog post.
MarkLogic is a multi-model NoSQL database that supports various data formats and interfaces. Datadog's integration with MarkLogic provides visibility into performance issues, allowing users to monitor storage performance, network activity, errors, and more. The integration helps maintain efficient data processing across large organizations like Airbus, the BBC, and the U.S. Department of Defense. By monitoring metrics such as query throughput, cache hit rates, XDQP traffic, App Server requests, and error logs, users can optimize their MarkLogic deployments and troubleshoot issues more effectively.
Nov 13, 2020
748 words in the original blog post.
CloudFront is a content delivery network (CDN) that minimizes latency by caching content on AWS edge locations worldwide. By configuring CloudFront to send real-time logs to Datadog, users can gain deep visibility into the health and performance of their distribution. This setup involves creating a real-time log configuration for an existing CloudFront distribution, specifying the fields included in the logs, and designating a Kinesis Data Stream as the endpoint for sending logs to Datadog. Additionally, users can route logs through Kinesis Data Firehose to Datadog, apply custom tags to their logs, analyze them using Log Explorer, create facets and measures based on log data, and set up log monitors to proactively notify teams of potential problems. This setup enables users to monitor their CloudFront distributions' performance, optimize cache hit ratios, track long-term trends in latency, and improve user experience by tuning parameters like cache size and time to live (TTL).
Nov 13, 2020
1,213 words in the original blog post.
Datadog's integration with MarkLogic provides a unified monitoring platform that helps customers optimize storage performance, detect network issues and connection failures, and debug database error messages. The integration allows users to track key metrics such as read query throughput, hit rates for caches, and XDQP throughput to identify potential performance bottlenecks. It also enables the detection of trends in error logs, enabling swift action to be taken when errors exceed expected levels. Additionally, the integration provides visibility into network activity, allowing users to detect traffic spikes and connection failures, and supports integrations with other storage technologies such as Hadoop, Amazon S3, and Azure Blob Storage.
Nov 13, 2020
759 words in the original blog post.
The CNCF's KubeCon North America 2020 is the premier event for adopters and technologists to learn about and work with the Kubernetes community. Datadog will be hosting several sessions, including "Datadog on Kubernetes Monitoring" and "Datadog on Autoscaling Applications Deployed to Kubernetes," which aim to provide insights into monitoring and autoscaling in Kubernetes environments. The company also plans to participate in two KubeCon sessions, covering topics such as PKI configuration mistakes and the stories of debugging complex issues with Kubernetes. Datadog will be hosting a virtual booth during the event, where attendees can learn about their latest platform additions and win prizes.
Nov 13, 2020
2,002 words in the original blog post.
Amazon Web Services (AWS) has been instrumental in driving the adoption of basic cloud computing services as well as the expansion and ease of use of managed infrastructure services. Over the years, AWS has expanded beyond basic compute resources to include tools like CloudWatch for AWS monitoring, managed infrastructure services like Amazon RDS for database management, and AWS Lambda for serverless computing. Datadog’s AWS integration aggregates metrics from across your entire AWS environment in one place and enables you to get full visibility into your highly dynamic services in order to efficiently investigate potential issues. This article explores key metrics that will help you monitor widely used services like Amazon EC2, EBS, ELB, RDS, ElastiCache, ECS, EKS, Lambda, and Fargate in full context with the rest of your infrastructure and applications.
Nov 12, 2020
5,157 words in the original blog post.
Yael Goldstein is discussing the importance of DNS in infrastructure and how complete visibility into internal and external DNS resolution is necessary to keep it healthy and performant. Datadog has announced new DNS monitoring features that provide a unified view of DNS traffic, enabling users to troubleshoot DNS end-to-end and ensure application performance and availability. The new Cloud Network Monitoring's DNS view provides insight into the health of internal DNS servers and service discovery, while Synthetic DNS tests help detect DNS failures and misconfigurations. With these features, users can assess the health of all internal DNS servers in a single view, investigate DNS communication from the client side, troubleshoot DNS server latency and failure, correlate DNS performance with server monitoring data, detect irregularities in DNS record mapping and resolution times, and identify potential problems before they affect users. The new feature also allows users to graph DNS errors by type and display visualizations of key DNS-specific health metrics, such as volume, response time, and failure rate of DNS requests. Additionally, Datadog provides integrations with popular DNS services like CoreDNS, PowerDNS, and Amazon Route 53, enabling users to correlate DNS flow data with performance metrics from across their entire environment.
Nov 10, 2020
1,594 words in the original blog post.
Datadog is now offering the ability to generate globally accurate aggregations of process metrics across any subset of applications and infrastructure, allowing users to store these metrics for up to 15 months. This feature enables users to analyze historical trends in their data, troubleshoot performance degradations, and detect future issues more proactively. With this new capability, users can also leverage percentile aggregates to spot outlying processes, apply unified service tagging, and use process metrics alongside other telemetry data to identify the root cause of issues. Additionally, Datadog's distribution metrics allow users to analyze high-cardinality data at a glance, providing non-percentile and percentile aggregations that concisely summarize resource consumption of processes running on any meaningful subset of hosts or containers. This feature enables users to pinpoint outliers, differentiate between primary and secondary workloads, and use process metrics in conjunction with other telemetry data to identify the root cause of issues. Furthermore, Datadog's alerts and SLOs enable teams to detect infrastructure issues proactively, allowing them to respond more effectively to incidents. Overall, this new feature provides users with a single place to view all their processes, enabling better visibility into their applications and infrastructure.
Nov 09, 2020
1,475 words in the original blog post.
The thread-per-core architecture is a design approach that aims to improve the efficiency and scalability of applications by utilizing modern hardware's capabilities. It involves running each core on a single thread, eliminating the need for threads and context switches. This approach can deliver significant performance gains, especially when combined with sharding, which allows multiple threads to work on different subsets of data. The Glommio library is an implementation of this architecture in Rust, designed to make it easy for developers to write efficient and scalable applications. It leverages Linux's io_uring API to manage I/O operations and provides a thread-per-core model that can run on Kubernetes infrastructure, with the potential to improve performance when matched with physical cores available in the underlying hardware.
Nov 02, 2020
2,843 words in the original blog post.