Home / Companies / Grafana Labs / Blog / August 2026

August 2026 Summaries

11 posts from Grafana Labs

Filter
Month: Year:
Post Summaries Back to Blog
Grafana Cloud’s Knowledge Graph introduces an instrumentation quality report that continuously evaluates how completely and correctly each service emits and connects telemetry such as metrics, logs, traces, profiles, service graph data, and Kubernetes metadata. The report uses focused automated checks to identify issues such as missing logs or traces, malformed service names, absent Kubernetes labels, and inadequate span metrics, then converts results into percentage-based quality tiers from Poor to Perfect. By emphasizing the connections among services, dependencies, infrastructure, and diagnostic signals rather than merely the volume of data collected, it aims to prevent investigation dead ends during incidents. Fleet-wide and service-level views let teams prioritize poorly instrumented services, inspect failing and passing checks, access documentation and validation queries, filter and export results, and rerun assessments on demand. The capability requires no added setup, updates as services are discovered or changed, and is also accessible through Grafana Assistant and the gcx CLI, helping organizations identify and resolve observability gaps before they affect incident response.
Aug 27, 2026 1,648 words in the original blog post.
Grafana Labs has open-sourced the Grafana AI SDK for Go, a shared framework created to reduce duplicated and inconsistent AI integration patterns across its Go-based backend teams. Inspired by Vercel AI SDK, it provides common Go interfaces for model calls, streaming, typed tools, structured output, multi-step agents, retries, fallbacks, and middleware, while supporting Anthropic, Amazon Bedrock, OpenAI Responses, and OpenAI-compatible APIs. The SDK is designed to work with Vercel’s AI SDK frontend protocol, enabling Go services to stream directly to React hooks such as useChat, useCompletion, and useObject without custom adapters. Its middleware includes Grafana-focused agent observability, logging, and Prometheus metrics, though developers remain responsible for security, authorization, validation, and data-handling decisions. Although the project is still evolving and does not claim complete parity with Vercel AI SDK, Grafana maintains public compatibility information and invites the Go and open-source communities to test, extend, and help shape it.
Aug 25, 2026 1,682 words in the original blog post.
Grafana Alloy can serve as a centralized Kubernetes telemetry gateway that receives metrics, logs, and traces from many teams, applies shared authentication, routing, buffering, protocol normalization, and cost-attribution labeling, then forwards data to Grafana Cloud. A Grafana Labs engagement sized a deployment for roughly 17 million active metric series, 1 TB per day each of logs and traces, and significant peak throughput, estimating about 195 GB of memory and 28 CPU cores spread across 33–35 small pods, with no CPU limits and an HPA configured for 30 to 100 replicas. The recommended architecture places Alloy behind an ingress controller while sending Alloy and cluster health metrics through an independent monitoring path directly to Grafana Cloud, avoiding loss of observability during gateway stress. Load testing with telemetrygen and customized k6 instances validated performance at and above anticipated production volumes while tracking accepted and refused data, exporter queues, resource use, and ingress errors. The resulting production fleet scaled to about 60 pods and handled higher-than-planned traffic without reported issues, while key lessons included budgeting for growth rather than baseline demand, configuring retries, controlling write-ahead-log memory growth with GOMEMLIMIT, maintaining adequate minimum replicas to absorb traffic bursts, and ensuring cluster node autoscaling can support pod expansion. Future options include queue-depth-driven autoscaling with KEDA, splitting deployments by telemetry signal, and using a message broker where stronger pipeline decoupling justifies its operational cost.
Aug 21, 2026 2,684 words in the original blog post.
Grafana 13.2 introduces enhancements aimed at making data exploration, query reuse, and dashboard management more accessible across teams. Saved queries, now generally available in Grafana Cloud and Enterprise, provide a shared, searchable library of vetted queries that can be created from dashboards, Explore, and annotations, reused with variable mapping, governed through role-based permissions, accessed through the command palette, and provisioned as code with Terraform. Grafana Assistant can help generate and refine queries in natural language across more than 30 data sources, while the new View panel sidebar, available in public preview across all editions, lets viewers investigate dense time-series panels through visualization controls and series fanout without editing dashboards. Other updates include expanded Git Sync support for GitHub Enterprise, GitLab and BitBucket webhooks, commit authorship information, a refreshed OSS and Enterprise homepage, Workload Identity Federation for BigQuery and Google Cloud Monitoring, a redesigned query-variable editor, and multi-panel grouping tools for dashboards.
Aug 20, 2026 1,501 words in the original blog post.
Grafana Labs evaluated whether its Grafana Cloud Knowledge Graph improves AI agents’ ability to diagnose production incidents compared with raw telemetry alone, finding that structured service, infrastructure, database, dependency, and health context can substantially improve root-cause analysis in certain multi-hop failures. In a replayed incident caused by excessive indexing of unique JSON values, agents with Knowledge Graph access identified the correct root cause in 15 of 16 trials, while telemetry-only agents succeeded once and frequently blamed a downstream query storm instead; the Knowledge Graph agents also used roughly half as many telemetry queries. The research highlights broader limitations of LLM-based debugging, including a tendency to pursue prominent but secondary signals, fabricate evidence when data is unavailable, and return inconsistent conclusions, costs, and investigation paths across identical runs. In a separate incident where the decisive evidence was a log line outside the Knowledge Graph, the graph neither improved nor worsened results, suggesting it can be useful without necessarily anchoring agents incorrectly. Production comparisons similarly indicated that Knowledge Graph-equipped agents generally used fewer tokens and, in a later study, completed investigations faster, although Grafana emphasizes that these are early findings from limited cases and that consistency, coverage of complex incident chains, and potential over-anchoring remain open questions.
Aug 19, 2026 2,397 words in the original blog post.
Grafana Cloud’s Synthetic Monitoring and Frontend Observability are presented as complementary tools for understanding frontend reliability: synthetic checks provide controlled, proactive signals that detect failures from known locations and scripted user journeys, while real user monitoring through the Faro SDK reveals the actual scope, conditions, and user experiences behind those signals. Synthetic monitoring alone can miss unanticipated user paths, browser or regional variations, and the number of users affected, whereas frontend observability captures real-session performance, JavaScript errors, contextual stack traces, and session replays. Combined workflows allow teams to investigate synthetic alerts by correlating failed checks with real-user sessions, quantifying blast radius, identifying shared factors such as browser, location, or error type, and making more informed incident-priority decisions. Real-user data can also expose unmonitored problem areas, guide the creation of new browser checks, establish performance thresholds based on observed p75 and p95 baselines, and help retire low-value tests. Because both data sources are available in Grafana Cloud, teams can instrument critical applications, compare signals by URL and time range, and build dashboards that connect uptime status with real-user error rates and traffic, creating a feedback loop intended to improve triage, alert relevance, test coverage, and collaboration between development and operations teams.
Aug 19, 2026 1,937 words in the original blog post.
Grafana Cloud Frontend Observability has introduced Session Replay in public preview, an add-on that visually reconstructs user interactions in web applications alongside existing metrics, logs, traces, and session timelines. Built on the Grafana Faro Web SDK, the feature helps teams investigate difficult or hard-to-reproduce frontend problems by showing how an interface appeared and changed during a user’s session, then linking visual events to technical data such as network requests and distributed traces. In a demo banking scenario, a replay revealed repeated failed payment attempts, HTTP 422 responses, and a backend trace indicating insufficient funds, while also exposing that a vague error message created user frustration. Privacy controls mask text and input values by default, offer configurable protections, restrict recording access through a dedicated RBAC role, and retain data for 30 days, although organizations remain responsible for consent and compliance. Session Replay requires Faro Web SDK version 2.8.2 or later, can be enabled by adding replay instrumentation, and is available through a public-preview waitlist as part of Grafana Cloud’s broader digital experience monitoring plans.
Aug 17, 2026 1,427 words in the original blog post.
k6 2.0 introduces k6 x docs, a command-line documentation tool intended to keep k6 API references, guides, examples, and best practices accessible directly within developers’ terminals and AI-assisted workflows. The command supports hierarchical topic lookups and keyword searches, caches documentation for offline use after the first request, and automatically serves content matching the installed k6 version. It also provides shell completions, terminal-friendly Markdown rendering, pager and width options, and an installable agent skill that helps AI coding assistants obtain authoritative k6 guidance instead of relying on stale information or web searches. Available by default in k6 2.0 and version 1.7.0 or later, the underlying Go documentation package can also be used by external tools such as the k6 MCP server, while the feature forms part of a broader set of AI-focused k6 x commands.
Aug 14, 2026 1,142 words in the original blog post.
Grafana Cloud’s new volumetric policy for Adaptive Traces automates dynamic trace sampling to preserve a more diverse and representative dataset within a specified storage budget. Unlike flat probabilistic sampling, which can allow high-volume services, endpoints, regions, or customers to dominate retained traces, volumetric sampling analyzes trace attributes, selects useful combinations with manageable cardinality, groups traces into buckets, and continually adjusts each bucket’s sampling rate as traffic changes. Grafana reports that, at the same sampling percentage, this approach produces roughly 25% higher information density than probabilistic sampling, retaining more unique information per byte stored. The policy complements Adaptive Traces features such as anomaly detection, diversity sampling, and explicit retention rules for latency, status, or audit-related traces; users can adopt it through onboarding or convert existing probabilistic policies with one click, while dropped traces remain retrievable for 24 hours.
Aug 13, 2026 1,661 words in the original blog post.
Grafana has introduced a Graphviz panel in private preview for Grafana 12.3 and later, enabling users to visualize workflows, business processes, service dependencies, network topologies, and decision trees alongside live observability data. Built on the open-source Graphviz project, the panel uses the DOT graph-description language to automatically arrange nodes and connections, avoiding the manual layout work required by canvas-based diagrams and replacing the discontinued FlowCharting plugin. It supports Builder, Code, and Query modes, allowing diagrams to be created visually, written directly in DOT, or generated dynamically from data sources such as service registries, Terraform state, or Kubernetes topology. Live query data can control node colors, labels, edge widths, tooltips, data links, and dashboard-variable-driven views, helping teams identify bottlenecks and unhealthy components in contexts such as payment pipelines, microservice maps, network weathermaps, incident runbooks, CI/CD pipelines, telemetry flows, and Shopify customer segmentation. Grafana Assistant can also generate DOT from natural-language requests, while the plugin’s automatic layout and live updates aim to make operational and business flows easier to understand from a shared dashboard.
Aug 11, 2026 2,771 words in the original blog post.
HCP Terraform and Terraform Enterprise self-hosted agents emit OpenTelemetry traces and metrics for run phases and write logs to standard output, enabling centralized monitoring when connected to Grafana Cloud through Grafana Alloy. The walkthrough uses Docker Compose to run an Alloy collector alongside an HCP Terraform agent, with the agent sending OTLP traces and metrics to Alloy while Alloy reads container logs through the Docker API, assigns them a matching service name, converts delta metrics to the cumulative format required by Grafana Cloud Metrics, and exports all signals to Grafana Cloud. It requires HCP Terraform agent execution mode, a configured agent pool and token, Grafana Cloud OTLP credentials, Terraform CLI, and Docker. After starting the stack and triggering Terraform runs in a workspace bound to the agent pool, users can inspect plan and apply traces in Tempo, agent metrics such as initialization and apply durations in Mimir, and execution logs in Loki. Key considerations include ensuring workspaces use agent rather than remote execution, running multiple licensed agents for concurrent workloads, separately collecting stdout and stderr logs because they are not emitted through OTLP, correlating logs with traces through a shared service name, and pinning collector versions for dependable production use.
Aug 05, 2026 1,665 words in the original blog post.