Home / Companies / Firebase / Blog / Post Details
Content Deep Dive

Eval-driven development: How we build better agent skills for Firebase

Blog post from Firebase

Post Details
Company
Date Published
Author
Charlotte Liang
Word Count
1,543
Company Posts That Month
2
Language
English
Hacker News Points
-
Post removed?
No
Summary

Firebase describes an evaluation-driven approach to improving AI coding agents’ use of its platform through Agent Skills, the Firebase CLI, and MCP servers. Skills provide focused, task-oriented instructions that complement official documentation by guiding agents toward current APIs, secure configurations, best practices, and tool usage, while MCP servers connect agents to Firebase tools and Google developer documentation. Firebase measures effectiveness with automated evaluations covering individual skill performance, skill activation accuracy, and end-to-end multi-product application deployment in real Firebase projects. Its development process begins by defining tests and baseline agent performance, then iteratively refining skills based on observed failures until results improve. In tests captured in July 2026, agents using Firebase Skills achieved a 78.0% pass rate compared with 31.7% without skills, while also using fewer input and output tokens and completing tasks more quickly. Firebase also uses persistent evaluation failures to identify usability problems in its CLI and MCP tools, leading to improvements such as clearer help text, smoother agent login, and fewer blocking interactive prompts.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
MCP 11 8,729 854 211 -20%
AI Agents 4 5,780 1,243 245 -15%
LLM 2 5,068 1,020 229 -34%
Developer Experience 1 462 233 85 -22%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.