How Uber evaluates AI agents at production scale
Blog post from Arize
Uber’s agent platform team argues that effective AI-agent evaluation depends less on adding tools than on embedding tracing, ownership, and feedback loops into everyday development. A production voice-agent incident, in which background speech about pizza caused a ride to be rerouted, revealed how offline tests can miss real-world failures; a spike in conversation length helped uncover the issue through production metrics. Uber therefore makes detailed tracing available from deployment, uses production behavior and agent context to generate evaluators and alerts, and continuously promotes reviewed failures into evolving offline datasets. The company also broadens evaluation beyond engineering by enabling product, design, and operations specialists to assess behavior based on their domain knowledge. Rather than treating evaluation as a launch threshold, Uber measures its value by whether it changes release decisions, product designs, datasets, or system architecture. Its longer-term vision is an eval copilot that uses traces, documentation, and prior results to identify recurring failures, recommend tests and agent changes, and help teams validate improvements while retaining human judgment over final decisions.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Observability | 10 | 3,175 | 737 | 186 | -24% |
| Platform Engineering | 6 | 1,191 | 259 | 79 | -17% |
| AI Agents | 5 | 5,780 | 1,243 | 245 | -15% |
| AI Coding Assistant | 2 | 1,513 | 470 | 139 | -19% |
| Harness engineering | 2 | 203 | 125 | 57 | -23% |
| AI Guardrails | 1 | 551 | 150 | 54 | +6% |
| LLM | 1 | 5,068 | 1,020 | 229 | -34% |
| Voice AI | 1 | 2,839 | 275 | 56 | -36% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.