September 2024 Summaries
24 posts from Grafana Labs
Filter
Month:
Year:
Post Summaries
Back to Blog
The text provides an overview of implementing automatic remediation workflows in Grafana Cloud to enhance system resilience and reliability during incidents, such as high CPU usage spikes during peak traffic periods. By setting up auto-remediation triggers, systems can automatically scale resources, restart failed services, reroute traffic, and rollback faulty deployments without human intervention, thus reducing downtime and maintaining service quality. The process involves using Grafana OnCall's escalation chains, configuring webhooks, and setting conditions for automatic responses to incidents. The article emphasizes the benefits of automation in reducing engineer workload, speeding up incident response, and minimizing human error, while also encouraging teams to experiment with low-risk automation solutions. Grafana Cloud offers a flexible platform with a free tier to help organizations start building these workflows effectively.
Sep 30, 2024
810 words in the original blog post.
The Business Suite for Grafana, developed by Volkov Labs and co-founded by Grafana Champion Daria Volkova, is a set of 10 versatile plugins designed to expand Grafana's functionality beyond observability into comprehensive web application development. These plugins offer a range of capabilities, such as file uploads, video streaming, web design enhancements, and more, catering to various business needs. Each plugin is tailored for specific tasks, including data input, media rendering, variable management, text formatting, news aggregation, satellite data access, chart creation, form integration, calendar management, and table organization. These plugins are maintained by Volkov Labs, ensuring ongoing support and compatibility with Grafana updates. The blog post outlines practical examples and tutorials for using these plugins, emphasizing their adaptability and potential to transform Grafana dashboards into fully functional web applications. The suite's impact is illustrated by its use in high-profile projects, such as JAXA's moon landing mission, showcasing its real-time monitoring capabilities. Volkov Labs remains committed to community engagement, continuously seeking feedback and inspiration to enhance their offerings in alignment with Grafana's evolving landscape.
Sep 27, 2024
2,342 words in the original blog post.
On September 26, 2024, Grafana Labs released several patch versions of their software to address a medium severity security vulnerability, CVE-2024-8118, found in Grafana Alerting's data source rule write endpoints. This vulnerability, stemming from an incorrect RBAC action check, could allow unauthorized users to create, edit, and delete alert rules, potentially leading to unauthorized data access. The issue affects versions from 8.5.0 to 11.2.0, and users are strongly advised to upgrade to the patched versions. Grafana Labs has already applied the necessary patches to Grafana Cloud and recommends administrators audit permissions to mitigate any risks. The vulnerability was initially introduced in March 2022, discovered and reclassified in August 2024, and subsequently fixed with a public release in September 2024. Users are encouraged to report any security vulnerabilities they find, and Grafana Labs commits to keeping reporters informed about the progress towards resolving such issues.
Sep 26, 2024
597 words in the original blog post.
Grafana Labs has released updates for Grafana Alloy and Grafana Agent to address high-severity security vulnerabilities, specifically CVE-2024-8975 and CVE-2024-8996, which could allow local users on Windows installations to escalate privileges. The issues stem from the installers not enclosing service executable paths in quotes, enabling a local user to execute unauthorized programs with elevated privileges. Grafana Alloy v1.4.1 and v1.3.4 and Grafana Agent v0.43.3 have been released with patches. Users are advised to perform clean installations rather than simple updates to fully address these vulnerabilities. The vulnerabilities were reported by a customer and subsequently fixed, with the timeline for the resolution provided, and Grafana Labs encourages users to report any additional security issues directly to them.
Sep 25, 2024
730 words in the original blog post.
Grafana Cloud has introduced several updates and features aimed at enhancing user experience and optimizing observability costs, following their ObservabilityCON 2024 event. Key updates include the renaming of the Explore apps suite to Drilldown apps for a queryless data exploration experience, the integration of Asserts technology for automated anomaly correlation, and the release of Adaptive Logs to minimize unnecessary log retention and reduce observability costs. Enhanced capabilities for synthetic monitoring and multi-cloud monitoring have been announced, alongside new data visualization options such as canvas actions, legend support in bar gauges, and improved cell inspection in tables. Updates also include streamlined onboarding for Grafana OnCall, a unified Slack integration for Grafana IRM, and UI improvements with an announcement banner for admins and navigation bookmarks. Additionally, new integrations and data source plugins have been introduced, including AI Observability and support for various platforms like Kafka, JVM, and Amazon Aurora, to facilitate comprehensive monitoring and data visualization.
Sep 25, 2024
1,728 words in the original blog post.
Grafana Labs has launched the Grafana Labs Startup Program to provide eligible startups with up to $100,000 in Grafana Cloud credits for 12 months or until their next funding round, aiming to alleviate financial pressures associated with scaling while offering flexible observability solutions. This initiative, announced at ObservabilityCON 2024, addresses the challenges startups face with cost-prohibitive observability solutions and proprietary lock-ins by offering an open and composable platform. New and existing Grafana Cloud users, especially those with less than $10 million in funding and fewer than 25 employees, are encouraged to apply. The program supports startups with managed services for logs, traces, metrics, and more, alongside a 20% discount post-credit depletion, access to enterprise plugins, 24/7 support, and collaboration with Grafana's integrations team. Telemetry management platform Datable.io, among the first participants, credits the program for enabling them to focus on business growth without worrying about monitoring costs, highlighting Grafana Cloud's flexibility and openness as key benefits.
Sep 24, 2024
768 words in the original blog post.
ObservabilityCON 2024 in New York showcased significant advancements by Grafana Labs aimed at enhancing observability through AI-driven features, intuitive tools, and comprehensive integrations. Key highlights included the introduction of Explore apps for a queryless data exploration experience and the integration of Asserts.ai technology into Grafana Cloud for streamlined anomaly correlation and troubleshooting. The Adaptive Telemetry features, such as Adaptive Logs and Metrics, promise cost efficiency by using AI/ML to manage data volume. The acquisition of TailCtrl is set to advance Adaptive Traces. Grafana Labs also launched a Startup Program offering cloud credits to encourage early adoption of their scalable solutions. Enhancements in AI/ML observability include tools for monitoring machine learning experiments with Nvidia and LLMs, along with new features in synthetic monitoring and multi-cloud observability for AWS, Azure, and Google Cloud. The conference emphasized efforts to simplify and enhance observability tools to make them more accessible and efficient for users.
Sep 24, 2024
1,312 words in the original blog post.
Adaptive Logs, a new feature in Grafana Cloud's Adaptive Telemetry suite, helps organizations reduce observability costs by optimizing log volumes and maintaining full system visibility. It leverages AI/ML techniques to analyze log data, identifying patterns and suggesting which low-value logs can be safely dropped without affecting insights. Adaptive Logs offers a user-friendly interface that provides daily recommendations based on real-time usage patterns, allowing users to easily manage their log data and cut down on unnecessary storage expenses. This feature is available for all Grafana Cloud tiers, including the free tier, and has already demonstrated significant cost savings, such as a reported 50% reduction in log volumes by early adopters like TeleTracking. By reducing noise and focusing on valuable logs, Adaptive Logs not only alleviates alert fatigue but also ensures that critical issues are more easily identified and addressed, enhancing the overall log management experience.
Sep 24, 2024
787 words in the original blog post.
Grafana Labs has announced its acquisition of TailCtrl, a startup specializing in adaptive trace sampling, to enhance the development of Adaptive Traces aimed at optimizing observability costs for complex IT architectures. This acquisition, revealed at ObservabilityCON 2024, aligns with Grafana Labs' strategy to improve visibility into distributed systems by addressing the challenge of managing and extracting value from the increasing volume of tracing data. TailCtrl, founded by Sean Porter, focuses on identifying and retaining the most valuable traces using tail-based sampling, thus transforming vast amounts of tracing data into actionable insights. The integration of TailCtrl's technology with Grafana's existing tools, such as Grafana Tempo, is expected to streamline anomaly detection, incident identification, and root cause analysis while supporting protocols like OpenTelemetry. This move supports Grafana Labs' broader vision of building a comprehensive suite of adaptive telemetry services, including Adaptive Metrics and Adaptive Logs, to enhance user efficiency and cost management.
Sep 24, 2024
567 words in the original blog post.
Grafana Cloud has enhanced its troubleshooting capabilities by integrating Asserts.ai, an AI/ML-driven tool, to automate the correlation of anomalies across application and infrastructure layers, thereby reducing the mean time to resolution (MTTR) for complex microservice-based applications. This integration, announced at ObservabilityCON 2024, introduces a suite of unified workflows that automate root cause analysis, making it easier even for junior engineers to diagnose issues effectively. The Asserts tool, designed for Prometheus and OpenTelemetry instrumentation, offers real-time monitoring and SLO-based alerts, providing a contextual layer to telemetry data and reducing alert fatigue. It features an RCA Workbench that helps visualize and prioritize system anomalies, and it integrates seamlessly with Grafana Cloud's observability solutions, enhancing the troubleshooting process. The integration is available to Grafana Cloud Advanced customers as of October 2, and includes the ability to perform domain-specific investigations through detailed dashboards and bi-directional workflows, offering a comprehensive and efficient approach to resolving application and infrastructure issues.
Sep 24, 2024
1,834 words in the original blog post.
Grafana Labs has introduced the Explore apps suite, which includes Explore Metrics, Explore Logs, Explore Traces, and Explore Profiles, to provide a queryless experience for extracting insights from observability data. These apps aim to simplify the interaction with metrics, logs, traces, and profiles by offering intuitive, point-and-click interfaces that eliminate the need for complex query languages, thereby making them accessible to both seasoned professionals and newcomers. The suite, designed to enhance usability across all core observability pillars, is now generally available in Grafana OSS, Grafana Enterprise, and Grafana Cloud, with the Explore Traces and Explore Profiles apps available in public preview. By enabling easy navigation through data, these tools help users quickly identify and address performance issues, visualize data trends, and optimize system operations, all without deep expertise in query languages. The transition to a more intuitive interface is part of Grafana's broader goal to transform how DevOps teams and SREs interact with their systems and data, promoting faster and more efficient problem-solving and system management.
Sep 24, 2024
1,122 words in the original blog post.
Companies are increasingly migrating to Grafana Cloud for its cost-effectiveness, improved mean time to resolution (MTTR), and enhanced user experience. Paradigm, a significant player in the cryptocurrency sector, moved to Grafana Cloud Logs to boost developer productivity and streamline issue diagnosis, resulting in greater engagement and trust in data. ComplyAdvantage, dealing with extensive microservices and data spans, found Grafana Cloud a better cultural fit due to its open-source nature and native integration with OpenTelemetry, which revolutionized their approach to data storytelling and business conversations. Actian, striving for a unified observability solution across its distributed IT infrastructure, achieved a single-pane-of-glass view and significant cost savings by consolidating multiple tools into Grafana Cloud, which also facilitated consistent observability across all development environments. These migrations highlight the benefits of Grafana Cloud in providing scalable, integrated observability solutions that align with the evolving needs of diverse organizations.
Sep 23, 2024
1,243 words in the original blog post.
The "Grafana's Big Tent" podcast episode discusses the importance of observability and automation in mitigating DDoS attacks, highlighting the roles of curiosity and proactive engineering in addressing such threats. Hosts Matt Toback and Dee Kitchen from Grafana Labs, along with Alex Forster from Cloudflare, emphasize the need for automation to handle attacks swiftly, reducing reliance on constant human monitoring. They explore the balance between logging necessary data for future insights and managing resource costs, advocating for a positive security model that focuses on understanding and adapting to what constitutes "normal" traffic patterns. The conversation underscores the value of curiosity in improving engineering practices, encouraging a proactive approach to system observability and security.
Sep 20, 2024
3,029 words in the original blog post.
Grafana Labs has announced the general availability of its Grafana OpenTelemetry distributions for Java and .NET, aiming to simplify the user experience and uphold open-source values in observability. The distributions address challenges new users face with OpenTelemetry's flexible yet complex components by offering tailored solutions that integrate seamlessly with Grafana Cloud Application Observability. Key motivations include enabling quick bug fixes, improving ease of use by addressing missing upstream features like service instance IDs, and enhancing cost efficiency by allowing users to discard unused metrics. Grafana faced development challenges, such as ensuring users are not locked into their distributions, and contributed to the OpenTelemetry specification to improve broader community use. Efforts to streamline authentication and reduce metric overload have been implemented to facilitate easier transitions between Grafana and upstream versions. Through these innovations, Grafana aims to better serve its users and the open-source community, encouraging adoption of its new distributions to enhance application performance monitoring.
Sep 19, 2024
807 words in the original blog post.
Doug Tidwell's blog post offers a comprehensive guide on using Grafana Cloud to monitor metrics and logs from Altinity.Cloud's ClickHouse clusters, showcasing the integration with Prometheus and Loki services. It details the process of setting up a remote Prometheus service within Grafana Cloud, connecting ClickHouse clusters via the Altinity Cloud Manager, and visualizing data transmitted to Prometheus. The post also covers configuring a hosted logs service through Grafana Loki, illustrating how to send log messages from ClickHouse clusters to an external Loki server and explore the log data within Grafana Cloud. The article underscores Grafana Cloud's "big tent" philosophy, allowing organizations to tailor their observability strategies by integrating various data sources into a unified platform. This integration is facilitated by Altinity's support for open-source real-time analytics, reinforcing their commitment to open-source software within the modern data stack.
Sep 19, 2024
2,049 words in the original blog post.
Grafana Labs invites participants to take part in their third annual Observability Survey, which aims to provide insight into the evolving observability landscape by highlighting community successes and challenges. The survey seeks to understand how many tools and data sources organizations are using, criteria for selecting observability tools, the role of AI in observability, and usage of Prometheus and OpenTelemetry, while ensuring respondent anonymity. The results, along with analysis, will be published early next year to offer context on industry trends and organizational practices. Participants are encouraged to share the survey within their networks and engage with Grafana Labs at upcoming events like ObservabilityCON, where they can complete the survey in person and receive swag.
Sep 18, 2024
335 words in the original blog post.
Martin Falch of CSS Electronics discusses how the integration of Amazon Athena with Grafana helps users visualize and analyze CAN bus data efficiently and cost-effectively. The CAN bus protocol, used for sensor data communication in vehicles and machinery, generates large volumes of data that need to be decoded and visualized for engineers involved in R&D, diagnostics, or predictive maintenance. The solution involves using Amazon Athena as a serverless and interactive analytics service within AWS, which, when combined with Grafana, allows users to self-deploy a data visualization workflow without coding. This setup, supported by AWS Lambda functions and Parquet data lakes, enables cost-effective analysis of vast data volumes, while retaining original timestamps for detailed insights. Furthermore, the structured data lake allows for multi-purpose data analysis, making it accessible via SQL interfaces for diverse applications, including Python, MATLAB, and Excel.
Sep 13, 2024
1,295 words in the original blog post.
OpenTelemetry and vendor neutrality: how to build an observability strategy with maximum flexibility
OpenTelemetry is an open-source project that emphasizes vendor neutrality, allowing users to avoid the pitfalls of being locked into proprietary systems by providing flexibility in telemetry collection and analysis. The project consists of three main layers: apps and infrastructure, telemetry collectors, and telemetry backends. OpenTelemetry's loose coupling of components enables users to mix and match different SDKs, APIs, and collectors, which facilitates interoperability and reduces dependency on any single vendor. While OpenTelemetry excels in decoupling telemetry collection from storage, it does not extend vendor neutrality to the backend layer, where data is stored and queried. This limitation arises because backend systems require specific customizations that can create vendor lock-in, similar to choosing a programming language for an application. Consequently, while OpenTelemetry's design allows for maximum flexibility at the telemetry collection stage, backend decisions remain more fixed. To maximize flexibility, users are advised to standardize on open protocols like OTLP and adopt a layered approach to implementation, enabling them to switch components without overhauling their entire system.
Sep 12, 2024
1,770 words in the original blog post.
Catchpoint has been introduced as an Enterprise data source for Grafana, enabling users to integrate and visualize Catchpoint's Digital Experience Monitoring (DEM) and Internet Performance Monitoring (IPM) capabilities within Grafana dashboards. This integration allows for real-time performance monitoring and enhanced data visualization, as users can create dynamic dashboards and conduct real-time analysis to quickly identify and resolve performance issues. The Catchpoint Enterprise data source plugin offers features like custom queries, comprehensive metrics, and alerting, providing a centralized, holistic view of digital experiences. It is available in Grafana Enterprise and Grafana Cloud, including a free-forever tier, with detailed documentation provided for installation and configuration. The community is encouraged to share their experiences and feedback to further enhance monitoring capabilities.
Sep 11, 2024
531 words in the original blog post.
Grafana's access management system can be streamlined by using teams to manage user permissions and access effectively, with the blog post explaining how to implement such a system by integrating with identity providers like Entra ID, Okta, and Keycloak. By adopting a team-based structure over organizations or custom roles, Grafana encourages collaboration and transparency while maintaining controlled isolation and data security. Users are guided through the setup process, involving synchronization of external user directories, mapping teams to resources, and configuring permissions for data sources and dashboards. The post highlights the benefits of using Terraform for consistent and scalable configuration management, ensuring version-controlled and reproducible setups. Future improvements in Grafana are anticipated to enhance resource management, team permissions, and identity provider integration. Grafana Cloud offers these features across all its plans, including a free tier, making it accessible for various use cases.
Sep 10, 2024
1,716 words in the original blog post.
Grafana has introduced a new history feature in Grafana Alerting, available in Grafana 11.2, aimed at enhancing root cause analysis by providing a comprehensive timeline of all state transitions for Grafana-managed alert rules. This feature allows users to track when specific alert instances began or stopped firing, aiding in system stability and outage prevention. The "History" page, accessible via the Grafana main menu, includes filters, an events chart, and an events table to help narrow down and analyze historical events based on labels or states. By displaying the frequency and timing of alert state changes, users can easily identify patterns and diagnose issues more effectively, improving incident response strategies. The feature is beneficial for DevOps engineers, system administrators, and SRE team members, and is available to all Grafana Cloud users, with additional resources and documentation provided for further learning and integration into alert management practices.
Sep 09, 2024
652 words in the original blog post.
Incident management in modern tech environments emphasizes proactive strategies, error budgets, and a culture of continuous improvement rather than blame. The "Grafana's Big Tent" podcast, featuring Grafana Labs team members and Alex Koehler from Prezi, explores these concepts, highlighting the importance of structured incident response systems and a culture that supports innovation and learning from mistakes. Prezi's approach, "you build it, you run it," aligns with Grafana Labs' practices, focusing on decentralized management and maintaining error budgets to balance risk and innovation. The discussion underscores the value of blameless post-incident reviews to foster a culture of accountability and improvement, and emphasizes the significance of centralization for managing infrastructure and tools like Grafana OnCall for efficient incident handling. The conversation also touches on the necessity of keeping systems updated and resilient through regular maintenance and automated processes, ensuring reliability and minimal disruption.
Sep 06, 2024
3,107 words in the original blog post.
Grafana Tempo 2.6 introduces significant performance improvements and a range of new features for TraceQL, built on the enhanced vParquet4 backend now set as the default block format. Key updates include the ability to discover span events by name, custom attributes, or elapsed time, and enhanced querying for span links and arrays, with native support for searching arrays like HTTP headers. The update removes RF3 metrics in favor of a more efficient RF1 implementation, which aims to improve performance, durability, and availability, leading to lower total cost of ownership and more efficient TraceQL searches. Enhanced features such as native histogram support, exemplars in TraceQL metrics, and improved memory consumption in blocklist polling are also part of this release, addressing the needs of multi-tenant cluster operators. The ongoing initiatives focus on making TraceQL metrics generally available and re-architecting to RF1 to meet performance objectives, inviting community engagement through forums and monthly calls. Users are encouraged to explore the latest updates on Grafana Cloud, which offers a free tier including 50GB of traces, logs, and 10K metric series.
Sep 05, 2024
862 words in the original blog post.
Grafana Labs has introduced new Enterprise data source plugins for Catchpoint, PagerDuty, and Amazon DynamoDB, allowing users to visualize data from these platforms within Grafana dashboards. These plugins are part of Grafana's commitment to offering flexible data access through its Enterprise and Cloud offerings, which aim to eliminate the need for users to maintain their own plugins. The Catchpoint plugin provides insights into digital experience monitoring through various query types, while the PagerDuty plugin allows for incident data visualization and management. The DynamoDB plugin enables interaction with Amazon's NoSQL database using PartiQL and supports time series visualization. Grafana has also launched a public roadmap to enhance community collaboration and transparency in plugin development, encouraging feedback and participation from users and developers to improve and expand its offerings. The roadmap aims to facilitate collaboration, provide early feedback, and promote upcoming plugins, with Grafana Cloud offering a free tier that includes access to these Enterprise data sources.
Sep 04, 2024
1,169 words in the original blog post.