Home / Companies / Arize / Blog / Post Details
Content Deep Dive

How Uber evaluates AI agents at production scale

Blog post from Arize

Post Details
Company
Date Published
Author
Sara Verdi
Word Count
2,645
Company Posts That Month
18
Language
English
Hacker News Points
-
Post removed?
No
Summary

Uber’s agent platform team argues that effective AI-agent evaluation depends less on adding tools than on embedding tracing, ownership, and feedback loops into everyday development. A production voice-agent incident, in which background speech about pizza caused a ride to be rerouted, revealed how offline tests can miss real-world failures; a spike in conversation length helped uncover the issue through production metrics. Uber therefore makes detailed tracing available from deployment, uses production behavior and agent context to generate evaluators and alerts, and continuously promotes reviewed failures into evolving offline datasets. The company also broadens evaluation beyond engineering by enabling product, design, and operations specialists to assess behavior based on their domain knowledge. Rather than treating evaluation as a launch threshold, Uber measures its value by whether it changes release decisions, product designs, datasets, or system architecture. Its longer-term vision is an eval copilot that uses traces, documentation, and prior results to identify recurring failures, recommend tests and agent changes, and help teams validate improvements while retaining human judgment over final decisions.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Observability 10 3,175 737 186 -24%
Platform Engineering 6 1,191 259 79 -17%
AI Agents 5 5,780 1,243 245 -15%
AI Coding Assistant 2 1,513 470 139 -19%
Harness engineering 2 203 125 57 -23%
AI Guardrails 1 551 150 54 +6%
LLM 1 5,068 1,020 229 -34%
Voice AI 1 2,839 275 56 -36%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.