Do evals the Airbnb way
Blog post from Pydantic
Airbnb’s generative AI evaluation approach recommends first reviewing roughly 100 real outputs and traces to identify recurring failures, then creating a small set of targeted evaluators using three complementary layers: deterministic programmatic checks for objective requirements, LLM judges for narrow interpretive questions, and human review for ground truth, disagreements, and judge calibration. The workflow demonstrated with Pydantic AI and Logfire instruments a policy-based support agent, validates structured outputs and tool usage, checks whether citations and escalation behavior follow explicit rules, and uses an LLM judge to assess whether responses are faithful to retrieved policy evidence. Evaluation results and traces can be inspected in Logfire to distinguish agent failures from flawed rubrics, while human annotations build a 50-to-100-example gold set used to measure and improve judge agreement. Confirmed failures become regression cases, and Logfire’s optimizer can propose prompt changes that reviewers validate against evidence and rerun against the same dataset. In production, inexpensive deterministic checks can run broadly while costlier LLM judging is sampled, with failed or uncertain cases feeding back into human review, calibration, offline experiments, and ongoing system improvements.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 4 | 1,189 | 251 | 109 | -83% |
| OpenTelemetry | 2 | 158 | 34 | 25 | -85% |
| Observability | 1 | 625 | 152 | 84 | -84% |
| Serverless | 1 | 149 | 44 | 30 | -80% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.