Introducing Supabase Evals
Blog post from Supabase
Supabase has open-sourced a tool called supabase/evals, designed to benchmark and assess the performance of AI agents, such as Claude Code and Codex, in handling real-world tasks within the Supabase ecosystem. This framework evaluates agents on tasks like building schemas and debugging Edge Functions, providing a scoring system that informs both a public benchmark and an internal regression suite. The initiative aims to measure the effectiveness of AI-driven development with Supabase, addressing the growing trend of using agents for project creation. By identifying areas where agents struggle, like inconsistent usage of declarative schemas and skill activation, Supabase seeks to refine its guidance and improve agent performance. The results are visualized on their website, offering insights into agent capabilities and areas needing improvement, with plans to enhance the tool further by expanding coverage and making scoring more stable.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| MCP | 4 | 35 | 6 | 6 | -100% |
| Edge Computing | 2 | 2 | 1 | 1 | -95% |
| AI Agents | 1 | 20 | 4 | 4 | -100% |
| LLM | 1 | 46 | 7 | 6 | -99% |
| Observability | 1 | 14 | 6 | 6 | -100% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.