Home / Companies / Supabase / Blog / Post Details
Content Deep Dive

Introducing Supabase Evals

Blog post from Supabase

Post Details
Company
Date Published
Author
-
Word Count
1,139
Company Posts That Month
1
Language
English
Hacker News Points
-
Post removed?
No
Summary

Supabase has open-sourced a tool called supabase/evals, designed to benchmark and assess the performance of AI agents, such as Claude Code and Codex, in handling real-world tasks within the Supabase ecosystem. This framework evaluates agents on tasks like building schemas and debugging Edge Functions, providing a scoring system that informs both a public benchmark and an internal regression suite. The initiative aims to measure the effectiveness of AI-driven development with Supabase, addressing the growing trend of using agents for project creation. By identifying areas where agents struggle, like inconsistent usage of declarative schemas and skill activation, Supabase seeks to refine its guidance and improve agent performance. The results are visualized on their website, offering insights into agent capabilities and areas needing improvement, with plans to enhance the tool further by expanding coverage and making scoring more stable.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
MCP 4 35 6 6 -100%
Edge Computing 2 2 1 1 -95%
AI Agents 1 20 4 4 -100%
LLM 1 46 7 6 -99%
Observability 1 14 6 6 -100%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.