Home / Companies / Pydantic / Blog / Post Details
Content Deep Dive

Do evals the Airbnb way

Blog post from Pydantic

Post Details
Company
Date Published
Author
-
Word Count
3,011
Company Posts That Month
4
Language
English
Hacker News Points
-
Post removed?
No
Summary

Airbnb’s generative AI evaluation approach recommends first reviewing roughly 100 real outputs and traces to identify recurring failures, then creating a small set of targeted evaluators using three complementary layers: deterministic programmatic checks for objective requirements, LLM judges for narrow interpretive questions, and human review for ground truth, disagreements, and judge calibration. The workflow demonstrated with Pydantic AI and Logfire instruments a policy-based support agent, validates structured outputs and tool usage, checks whether citations and escalation behavior follow explicit rules, and uses an LLM judge to assess whether responses are faithful to retrieved policy evidence. Evaluation results and traces can be inspected in Logfire to distinguish agent failures from flawed rubrics, while human annotations build a 50-to-100-example gold set used to measure and improve judge agreement. Confirmed failures become regression cases, and Logfire’s optimizer can propose prompt changes that reviewers validate against evidence and rerun against the same dataset. In production, inexpensive deterministic checks can run broadly while costlier LLM judging is sampled, with failed or uncertain cases feeding back into human review, calibration, offline experiments, and ongoing system improvements.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 4 1,189 251 109 -83%
OpenTelemetry 2 158 34 25 -85%
Observability 1 625 152 84 -84%
Serverless 1 149 44 30 -80%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.