How we chose the model behind Topics with Baseten
Blog post from Braintrust
Topics utilizes a small, cost-effective model to achieve active observability by reading and summarizing production traces with a large language model (LLM) and clustering the summaries, allowing users to monitor what their agents are doing without manually reviewing logs. The development and optimization of this model involved collaboration between Baseten and Braintrust, focusing on balancing affordability and quality, with iterations primarily targeting the summarization step. An off-the-shelf Gemma 4B model was initially used, and through prompt adjustments and benchmarking, the model was refined to improve label correctness and recall of issues at a fraction of the cost of frontier models. The production setup achieved 82.2% label correctness and demonstrated that a well-crafted small model, paired with accurate prompts and examples, can outperform more expensive models on specific metrics. This approach to model optimization and active observability is not unique to Topics and can be applied to other products that rely on LLMs and large datasets.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 3 | 6,942 | 1,215 | 234 | +11% |
| Observability | 2 | 3,732 | 711 | 187 | -12% |
| AI Model Fine-tuning | 1 | 887 | 199 | 73 | +20% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.