Home / Companies / Braintrust / Blog / Post Details
Content Deep Dive

Eval feedback loops

Blog post from Braintrust

Post Details
Company
Date Published
Author
Ankur Goyal
Word Count
1,002
Company Posts That Month
3
Language
English
Hacker News Points
-
Post removed?
No
Summary

In the field of AI engineering, establishing a set of real-world examples, known as "evals", is crucial to understand how changes will impact end users. However, finding great eval data, identifying interesting cases in production, and tying user feedback to evals can be challenging problems. To overcome these issues, connecting real-world log data to evals allows for the evaluation of new and interesting cases in the wild, enabling improvements and avoiding regression. This is achieved by structuring evals as a function of data, prompts/code, and scoring functions, and utilizing tools like Braintrust's Eval function that streamlines this process. By capturing and utilizing logs, teams can power their evals with real-world examples, making it easier to identify interesting cases and improve AI products. As teams scale, filtering logs to only consider the most interesting ones becomes critical, and using filters, tracking user feedback, or running online scores can help uncover test cases that need improvement. Braintrust's solution provides a unified UI for exploring logs and evals, automating code reuse, and storing datasets in a cloud environment.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 2 3,398 379 136 +44%
Observability 1 1,227 261 93 -15%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.