Home / Companies / Grafana Labs / Blog / January 2023

January 2023 Summaries

21 posts from Grafana Labs

Filter
Month: Year:
Post Summaries Back to Blog
Kubernetes application performance monitoring (APM) is essential for maintaining application health and ensuring a positive user experience, but it presents challenges as Kubernetes does not natively provide easy application monitoring. APM focuses on monitoring individual applications within a Kubernetes cluster, while Kubernetes monitoring targets the performance of the cluster itself. Key metrics for APM include request rate, response time, error rate, memory usage, CPU usage, persistent storage usage, and uptime. To monitor applications, developers can build metrics logic into applications, use Kubernetes sidecar containers, or deploy Grafana Agent, with each method offering distinct advantages and potential drawbacks. Effective monitoring involves determining how to collect and visualize metrics and deciding on alerting strategies. Grafana provides tools and integrations to facilitate Kubernetes APM, helping improve both application quality and Kubernetes workflows.
Jan 31, 2023 2,314 words in the original blog post.
Grafana Labs webinars offer a platform for users to learn about the Grafana Stack and engage with the open-source community through live demonstrations, discussions on current releases, and insights into new products. These webinars cater to both beginners and advanced users of the Grafana LGTM stack—comprising Loki, Grafana, Tempo, and Mimir—and cover topics such as reducing total cost of ownership (TCO) and managing high cardinality metrics. Participants can benefit from live Q&A sessions with Grafana Labs experts to address specific observability challenges. Webinars are available live across various time zones and on-demand for flexible viewing. The upcoming sessions include topics like reducing mean time to resolution (MTTR), designing effective dashboards, and utilizing Grafana plugins for data visualization across different platforms, such as Microsoft Azure and Google Cloud.
Jan 30, 2023 621 words in the original blog post.
Distributed tracing is an essential tool for monitoring Kubernetes applications, providing insights into application health and performance by tracking interactions between microservices. Based on Google's Dapper paper, it uses spans—unique units of work with identifiers—to reconstruct transactions and understand service interactions, offering a more comprehensive view than traditional logging. Distributed tracing is crucial for Kubernetes observability, allowing developers to identify service dependencies, detect anomalies, and enhance debugging through end-to-end visibility. Tools like Grafana Tempo, supporting protocols such as OpenTelemetry, Zipkin, and Jaeger, facilitate distributed tracing by collecting and storing telemetry data efficiently, offering scalable and cost-effective solutions for tracing in complex cloud environments. OpenTelemetry, the convergence of OpenTracing and OpenCensus, is emerging as a standard, backed by major IT vendors, and offers robust instrumentation support across multiple programming languages, enhancing the capability of distributed tracing tools like Grafana Tempo to integrate seamlessly without modifying application code.
Jan 27, 2023 1,885 words in the original blog post.
FOSDEM 2023, the world's largest open-source conference, took place in Brussels with around 8,000 attendees and featured 759 lectures and 55 developer rooms. Grafana Labs had a presence with a dedicated stand for the fifth year, offering opportunities to interact with users and discuss monitoring and observability solutions. The event highlighted OpenTelemetry, continuous profiling, and eBPF as key focus areas, with various talks and devroom sessions organized by Grafana contributors. Attendees could engage directly with the Grafana team, share feedback, and explore the integration of Grafana with diverse data sources, exemplified by unique use cases such as monitoring bee hives.
Jan 26, 2023 515 words in the original blog post.
Kubernetes monitoring is a complex process that requires attention to various metrics across different layers, including services, containers, pods, deployments, nodes, and clusters, to ensure the health and performance of a project. Key tools such as kube-state-metrics, Metrics Server, and cAdvisor provide insights into resource utilization and the state of Kubernetes objects, which can be crucial for identifying and resolving potential issues like memory leaks. While cluster-level metrics offer a high-level overview, more granular monitoring of nodes, pods, and containers is essential for pinpointing performance bottlenecks and optimizing resources. Prometheus and Grafana are recommended for scraping, analyzing, and visualizing these metrics, allowing for effective monitoring and management of Kubernetes environments. Kubernetes Monitoring in Grafana Cloud offers a comprehensive solution with prebuilt dashboards and alerts, accessible to all users, including those on the free tier.
Jan 25, 2023 1,648 words in the original blog post.
Grafana has released versions 9.3.4 and 9.2.10 to address several security vulnerabilities, including CVE-2022-23552, CVE-2022-41912, and CVE-2022-39324. CVE-2022-23552 involves a stored XSS vulnerability in the Geomap and Canvas plugins, which can be exploited by users with Editor roles to execute arbitrary JavaScript in dashboards. CVE-2022-41912 pertains to a SAML privilege escalation issue in Grafana Enterprise, where unsigned assertions in SAML responses could be misinterpreted as signed, potentially allowing unauthorized access. CVE-2022-39324 involves the spoofing of the originalUrl parameter in snapshot functionality, which could mislead users by redirecting them to malicious URLs. Patches have been applied to Grafana Cloud, and users are advised to upgrade their instances to mitigate these risks. Security vulnerabilities can be reported to Grafana Labs via an encrypted message using their PGP key.
Jan 25, 2023 711 words in the original blog post.
Katrina Turner, a software engineer at the Energy Sciences Network (ESnet), highlights the crucial role of Grafana in visualizing network data to ensure the smooth operation of ESnet, a high-performance network funded by the U.S. Department of Energy. By leveraging Grafana's extensible architecture, ESnet has developed custom visualizations such as the Sankey Panel, Slope Graph, Bump Chart, Chord Diagram, and Matrix Panel to monitor network activity and enhance their data analysis capabilities. These visualizations, integral to projects like Stardust and NetSage, help predict capacity needs and visualize network flow, connectivity, and user rankings over time. The collaboration and contributions to the Grafana community illustrate the evolving capabilities of Grafana, especially after a transition to React-based plugins, allowing for tailored solutions that meet specific analytical needs of the network.
Jan 24, 2023 1,106 words in the original blog post.
Kubernetes, a leading container orchestration system, enables management and deployment of containers across diverse environments and requires real-time data on cluster activities through Kubernetes events for effective alerting and monitoring. These events, generated by state or configuration changes in cluster resources like pods and nodes, provide critical insights but are ephemeral, lasting only an hour without a retention mechanism. Integrating tools like Grafana with Kubernetes can enhance observability by collecting and visualizing this event data, allowing for the detection and timely resolution of issues. Grafana, part of the CNCF ecosystem, offers solutions for managing Kubernetes events through its lightweight Grafana Agent, which forwards event data to the Grafana LGTM Stack, enabling users to create dynamic dashboards, set specific alerts, and aggregate events over time for comprehensive monitoring. This approach helps maintain the health of cloud-native applications by ensuring that problems are identified and addressed efficiently, with options for both self-hosted and cloud-based Grafana instances available for different user needs.
Jan 23, 2023 1,998 words in the original blog post.
Kubernetes has gained popularity among DevOps engineers for deploying and managing containerized applications, with monitoring being a critical component for maintaining optimal performance. Prometheus, an open-source monitoring tool, is commonly utilized for monitoring Kubernetes clusters due to its ability to collect and store metrics as time series data. The Prometheus Operator simplifies the deployment and management of Prometheus instances within Kubernetes by abstracting complex configurations and supporting dynamic resource updates. Kubernetes operators enhance the Kubernetes ecosystem by automating application deployment, management, and monitoring through custom resources and controllers, thus reducing manual intervention and improving efficiency. The integration of Grafana for data visualization further aids in analyzing the health of clusters. Kubernetes Monitoring in Grafana Cloud offers a comprehensive solution for monitoring Kubernetes infrastructures, providing users with access to metrics, logs, and prebuilt dashboards.
Jan 19, 2023 2,323 words in the original blog post.
Grafana Labs has developed a recruitment data dashboard using their own Grafana tool to tackle challenges in visualizing recruitment metrics and making data-driven decisions amidst ambitious hiring goals. This dashboard consolidates data from various sources, including manual trackers, Greenhouse Recruiting, and Google BigQuery, into a comprehensive platform that enables stakeholders to easily access key recruitment insights such as hires per recruiter, time-to-hire, and acceptance rates. The development involved collaboration with data analytics and IT teams to ensure secure data handling and effective visualization. The dashboard, which has improved hiring process efficiency by identifying trends and areas for improvement, is planned to be open-sourced, allowing other organizations to leverage its framework. Future enhancements include integrating machine learning and AI for predictive analytics and possibly introducing job planning scorecards.
Jan 18, 2023 1,147 words in the original blog post.
ObservabilityCON on the Road 2023 is a series of live events organized by Grafana Labs to bring the open-source observability community together across the U.S., Europe, and the Asia-Pacific region. Following the success of last year's in-person events, this initiative aims to share the latest trends in open-source observability and foster connections among community members utilizing the Grafana LGTM Stack, which includes Loki for logs, Grafana for visualization, Tempo for traces, and Mimir for metrics. Each event will feature hands-on demos, technical deep dives, Q&A sessions, and an "Ask the Experts" booth for addressing technical queries on Prometheus, OpenTelemetry, and related topics. Registration and detailed agendas are now available for cities including Berlin, the Bay Area, Amsterdam, Chicago, Singapore, and Sydney, offering a platform for participants to discuss their diverse applications of Grafana, from home projects to significant endeavors like the NASA Astra launch.
Jan 18, 2023 390 words in the original blog post.
At Adobe, the integration of OpenTelemetry, Grafana, Grafana Mimir, and Grafana Tempo into their CI/CD pipeline has streamlined their observability operations, making them effectively invisible yet impactful for developers. By consolidating their tools from around 20 to just four, Adobe has embraced a flexible, open observability strategy that enhances developer productivity and provides a consistent experience. This approach enables developers to quickly deploy Grafana dashboards and monitor performance without specialized assistance, even across Adobe's three clouds and 26,000 employees. The strategy is driven by three key principles: leveraging open standards, focusing on productivity, and creating a consistent experience. With prebuilt dashboards available within minutes, developers can visualize and analyze metrics and traces efficiently, allowing them to concentrate more on feature development rather than observability.
Jan 13, 2023 788 words in the original blog post.
Grafana Labs addressed a security update from CircleCI, urging all former users to rotate secrets due to potential vulnerabilities. Although Grafana Labs no longer uses CircleCI, they proactively rotated or invalidated any previously used secrets and conducted a thorough review, finding no signs of suspicious activity or compromise. As a result, they identified two GPG keys used for signing binaries and Helm charts, prompting the rotation of these keys, especially for users who installed Grafana through their package repositories or used binary releases. Helm chart signing is discontinued due to its limited utility and associated risks. Updated instructions are provided for Debian/Ubuntu and rpm-based systems to replace the old GPG keys with new ones, while Helm chart users are advised to remove deprecated keys. Grafana Labs also invites users to report any security vulnerabilities via encrypted emails using their PGP key, and they maintain a dedicated blog category for security announcements and updates.
Jan 12, 2023 1,159 words in the original blog post.
Azure Managed Grafana, a service developed through a partnership between Grafana Labs and Microsoft, enables Azure customers to deploy secure, scalable Grafana instances for visualizing and analyzing data from various sources within the Azure cloud platform. Since its general availability in August 2022, it has become a popular choice for operational dashboards, with over a million active installations. The service now offers an upgrade path to Grafana Enterprise, which enhances IT investments by integrating data through plugins for a range of services like ServiceNow, Splunk, and Snowflake, while providing comprehensive support and professional services. Grafana Labs emphasizes interoperability and composability in observability, offering a complete stack with Grafana Cloud and on-premise solutions for logs, metrics, and traces, catering to diverse customer needs. This strategic expansion aims to meet users' requirements wherever they are, driven by the vibrant Grafana community.
Jan 11, 2023 355 words in the original blog post.
Grafana Loki is a scalable and multi-tenant log aggregation system that simplifies log queries by using labels instead of indexing log content, making it cost-effective and easy to operate for DevOps and SRE teams. To enhance query performance in Loki, it is recommended to reduce and refine searches with five key strategies: utilizing label or log stream selectors to narrow down log volumes, employing line filters to swiftly discard irrelevant data, selecting appropriate time ranges to minimize data processing, using LogQL parsers to extract structured fields efficiently, and implementing recording rules for frequently used metrics to boost efficiency in long-range queries. These practices optimize Loki's performance, enabling faster insights from log data while maintaining cost-effectiveness. Grafana Cloud offers an accessible platform to get started with these capabilities, providing a free tier and various plans to suit different needs.
Jan 10, 2023 897 words in the original blog post.
Grafana Machine Learning has introduced a new feature that enhances its forecasting capabilities by accounting for holidays, addressing the limitations of its existing model, Prophet, which already considers yearly, weekly, and daily seasonality in time series data. This feature allows users to incorporate specific predictable shifts in data by informing the model of past and future holiday occurrences, either by directly adding them in the UI or by providing a public iCalendar address. The functionality proves particularly useful for adjusting forecasts during holidays, as demonstrated with U.S. public holidays, where standard models might misinterpret traffic patterns as regular weekdays. By linking holidays to forecasts, Grafana Machine Learning can more accurately predict behaviors during these periods, thus improving alert accuracy and reducing false positives. This advancement is part of Grafana Cloud Pro and Advanced plans, offering customers the ability to refine their time series forecasts by integrating real-world events into their data analysis.
Jan 09, 2023 1,122 words in the original blog post.
The post discusses the integration of the k6 browser module, an experimental feature of Grafana k6, which adds browser-level APIs to enhance web performance testing by allowing interactions with browsers to collect comprehensive metrics. This module enables a hybrid approach by combining protocol-level and browser-level tests in a single script, offering a more realistic end-to-end assessment of user experiences by simulating interactions through the browser. The k6 browser module is built to be compatible with the Playwright API, making it easier for users familiar with NodeJS to adopt it, and it supports both synchronous and asynchronous operations while providing detailed performance insights. The hybrid testing strategy addresses the limitations of traditional load testing methods by providing a unified view of frontend and backend performance, which helps identify issues that may not be apparent when testing solely at the protocol level. The article encourages community involvement for feedback and further development, emphasizing that browser automation is a crucial component of web application testing.
Jan 08, 2023 2,224 words in the original blog post.
David Calvert, a site reliability engineer at Powder, describes his experience in building and operating Kubernetes clusters using Grafana and Prometheus for monitoring purposes. At Powder, an AI-powered gaming platform, the team uses Amazon EKS and other cloud services to manage their applications, relying on kube-prometheus-stack to deploy Prometheus and Grafana for metrics visualization. Calvert developed his own open-source Grafana dashboards for Kubernetes monitoring, which gained wider attention after he promoted them through a blog post. These dashboards, part of the dotdc/grafana-dashboards-kubernetes project, have been enhanced by community contributions and are now used to monitor and optimize both applications and infrastructure at Powder. The team is also transitioning to the Grafana LGTM Stack by adopting Grafana Tempo and OpenTelemetry, aiming for better integration of observability tools. Grafana offers a full Kubernetes Monitoring solution in Grafana Cloud, providing out-of-the-box access to metrics, logs, and events for Kubernetes users.
Jan 06, 2023 994 words in the original blog post.
The tutorial explains how to use the Grafana Ansible collection to deploy and manage Grafana Agents across multiple Linux hosts, simplifying the process of monitoring numerous machines. By incorporating the grafana_agent Ansible role, users can deploy Grafana Agents on eight Linux hosts simultaneously and monitor them via Grafana Cloud, utilizing prebuilt dashboards from the Linux Node integration for enhanced visibility. The setup involves configuring SSH access, creating an Ansible inventory, and writing an agent configuration file, followed by deploying the agents using an Ansible playbook. Users can verify the successful ingestion of logs and metrics into Grafana Cloud using the Explore feature, enabling efficient monitoring through customizable dashboards. The guide emphasizes the scalability of this approach, highlighting that the process can be expanded to manage more hosts by updating the inventory file and re-running the playbook.
Jan 05, 2023 1,202 words in the original blog post.
In 2017, Just Eat Takeaway.com (JET) faced the challenge of scaling its real-time monitoring system due to its rapid growth from a startup to a global scaleup, which led to issues with data accuracy. To tackle these challenges, JET re-architected its operations by adopting a hybrid model that combines Grafana Cloud with local Graphite infrastructure, enabling them to handle 4 billion logs and 120 terabytes of data daily. This approach allowed them to improve reliability and resilience, reduce complexity, and provide a unified interface for telemetry data while facilitating innovation and growth. By leveraging Grafana Cloud, the company has successfully integrated structured logging and enriched contextual data, enhancing their observability platform's efficiency and consistency. Looking forward, JET plans to focus on OpenTelemetry to provide a vendor-neutral observability language, which they anticipate will improve performance and free up engineering resources, fostering an innovative environment.
Jan 04, 2023 739 words in the original blog post.
JPMorgan Chase has developed a proactive monitoring tool using Grafana and Prometheus to enhance the stability and efficiency of their trading platform. Faced with market volatility and the challenges of remote work due to the pandemic, the team sought a solution that integrated their existing systems with modern tools to provide real-time insights into system capacity and potential issues. Their custom tool, TradeMon, leverages Grafana's visualization capabilities to offer transparent and shareable reports, facilitating better decision-making from the C-suite to the trading desk. Additionally, they are collaborating with their AIOps team to incorporate AI and machine learning for dynamic error modeling and to automatically identify new service level objectives and market anomalies. This approach aims to further improve system reliability and responsiveness, with plans for expanded AI integration and enhanced reporting tools.
Jan 04, 2023 621 words in the original blog post.