Eval-driven development: How we build better agent skills for Firebase
Blog post from Firebase
Firebase describes an evaluation-driven approach to improving AI coding agents’ use of its platform through Agent Skills, the Firebase CLI, and MCP servers. Skills provide focused, task-oriented instructions that complement official documentation by guiding agents toward current APIs, secure configurations, best practices, and tool usage, while MCP servers connect agents to Firebase tools and Google developer documentation. Firebase measures effectiveness with automated evaluations covering individual skill performance, skill activation accuracy, and end-to-end multi-product application deployment in real Firebase projects. Its development process begins by defining tests and baseline agent performance, then iteratively refining skills based on observed failures until results improve. In tests captured in July 2026, agents using Firebase Skills achieved a 78.0% pass rate compared with 31.7% without skills, while also using fewer input and output tokens and completing tasks more quickly. Firebase also uses persistent evaluation failures to identify usability problems in its CLI and MCP tools, leading to improvements such as clearer help text, smoother agent login, and fewer blocking interactive prompts.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| MCP | 11 | 8,729 | 854 | 211 | -20% |
| AI Agents | 4 | 5,780 | 1,243 | 245 | -15% |
| LLM | 2 | 5,068 | 1,020 | 229 | -34% |
| Developer Experience | 1 | 462 | 233 | 85 | -22% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.