Home / Companies / Braintrust / Blog / Post Details
Content Deep Dive

How to set up quality, cost, and latency alerts for AI agents

Blog post from Braintrust

Post Details
Company
Date Published
Author
Braintrust Team
Word Count
3,538
Company Posts That Month
6
Language
English
Hacker News Points
-
Post removed?
No
Summary

AI agents require monitoring beyond conventional infrastructure signals because they can remain available while output quality declines, costs rise, or latency worsens due to model changes, prompt revisions, retries, tool failures, or growing context. The guidance recommends instrumenting complete agent traces with consistent metadata for environments, models, prompts, agents, tools, users, workloads, token usage, costs, timing, errors, and retries, while using child spans to capture model calls and orchestration steps. Quality should be measured through asynchronous online scoring, combining deterministic checks and LLM-based evaluators at span, trace, or group scope, with sampling rates chosen to provide reliable production signal. Braintrust alerts can use per-log SQL conditions for individual failures such as low scores, expensive runs, slow requests, or errors, and Time window alerts for aggregate measures such as percentiles, rates, and average scores. Alerts should be segmented by meaningful operational dimensions, calibrated from historical baselines, routed to responsible teams through Slack or webhooks, tested and tuned using real production behavior, and assigned clear ownership to limit alert fatigue. Confirmed, reproducible incidents should be converted into evaluation cases and included in CI to prevent future regressions in quality, cost, latency, or agent behavior.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 9 747 162 79 -85%
AI Agents 6 931 231 103 -84%
Observability 4 472 102 54 -85%
Harness engineering 1 33 23 14 -84%
Real-time 1 649 155 80 -85%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.