OpenTelemetry for LLM tracing: a guide to instrumenting agents and routing spans anywhere
Blog post from Braintrust
Traditional Application Performance Monitoring (APM) tools like Datadog, Grafana, and Honeycomb often fall short in effectively monitoring Large Language Model (LLM) applications, as they focus primarily on metrics such as latency, error rates, and system health, which may not reflect the quality of the model's output. OpenTelemetry offers a solution by integrating structured telemetry at the LLM layer, capturing detailed data about prompts, retrievals, tool calls, and model responses, which standard APM tools typically overlook. By implementing OpenTelemetry's GenAI semantic conventions, organizations can trace and evaluate LLM applications more effectively, ensuring output quality and system observability are interconnected. This approach allows teams to route spans to multiple backends, such as Braintrust for output scoring and traditional APM tools for operational monitoring, without needing to modify existing telemetry paths. Through distributed tracing and the use of both automatic and manual spans, OpenTelemetry provides a comprehensive view of LLM workflows, enabling teams to debug, assess quality, and turn production failures into test cases, thereby maintaining robust and reliable LLM application pipelines.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| OpenTelemetry | 36 | 965 | 147 | 50 | 0% |
| Observability | 27 | 3,732 | 711 | 187 | -12% |
| LLM | 22 | 6,942 | 1,215 | 234 | +11% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.