Home / Companies / OpenRouter / Blog / Post Details
Content Deep Dive

How to Gate Pull Requests on LLM Evals in CI

Blog post from OpenRouter

Post Details
Company
Date Published
Author
OpenRouter
Word Count
5,612
Company Posts That Month
3
Language
English
Hacker News Points
-
Post removed?
No
Summary

LLM evaluation gates can protect pull requests from prompt or agent regressions that ordinary CI tests miss by keeping a fixed, reviewed set of test cases in the repository, running them when relevant files change, and blocking merges when results fall below a measured threshold. The guide demonstrates a support-agent example using an OpenRouter-backed Node script that checks required and forbidden response strings, samples each case multiple times, uses majority voting to reduce nondeterministic failures, pins requests to one provider, distinguishes evaluation infrastructure failures from quality regressions through exit codes, and reports runtime and model cost. It recommends job-level rather than workflow-level path filtering in GitHub Actions, requiring both the change-detection and evaluation jobs in branch protection, protecting prompts and eval sets with CODEOWNERS, and calibrating thresholds from repeated unchanged-branch runs rather than choosing arbitrary targets. For tool-calling agents, Ori Eval can test actions such as whether the correct tool was called and can use calibrated LLM judges, while platforms including DeepEval, Braintrust, Arize, and Galileo offer metrics, dashboards, and run history at varying levels of dependency. Key operational concerns include assertion quality, model parameter support, provider variation, cost, speed, secret handling for forked pull requests, and ensuring that skipped or failed filter jobs cannot allow untested changes to merge.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 12 No monthly metrics for this publish month.
Secrets Management 5 No monthly metrics for this publish month.
AI Agents 1 No monthly metrics for this publish month.
Observability 1 No monthly metrics for this publish month.
Real-time 1 No monthly metrics for this publish month.
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.