Home / Companies / Braintrust / Blog / Post Details
Content Deep Dive

How Dropbox automates evals for conversational AI

Blog post from Braintrust

Post Details
Company
Date Published
Author
Ornella Altunyan
Word Count
1,544
Company Posts That Month
18
Language
English
Hacker News Points
-
Post removed?
No
Summary

Dropbox, a prominent cloud storage and collaboration platform, developed Dropbox Dash, an AI-powered tool designed for universal search and organization across connected applications, highlighting the significance of AI evaluation alongside model training in the foundation-model era. The development of Dash involved creating a structured evaluation framework that approaches experiments with the same rigor as production code, shifting from ad-hoc testing to systematic evaluation. This process involved curating diverse datasets, including both public sources like Google's Natural Questions and internal datasets from Dropbox employee usage to mirror real-world complexity. Dropbox utilized large language models (LLMs) as judges to assess factual correctness, citation, and formatting, moving beyond traditional metrics such as BLEU and ROUGE. They adopted Braintrust as an evaluation platform to manage datasets and experiments, ensuring reproducibility and tracing regressions through defined metrics and automated checks. By automating evaluation in the development-to-production pipeline, Dropbox reduced the risk of regressions, integrating continuous improvement by mining low-scoring outputs for new dataset iterations. This approach emphasized the importance of versioning datasets, calibrating model judges, and treating prompt changes as code changes, transforming their AI development process into one that ensures reliable and trustworthy AI products.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 5 4,863 783 205 +34%
Voice AI 3 971 139 44 +45%
AI Guardrails 1 285 103 50 -30%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.