Home / Companies / OpenRouter / Blog / Post Details
Content Deep Dive

LLM-as-a-Judge: Score AI Agent Outputs Automatically

Blog post from OpenRouter

Post Details
Company
Date Published
Author
OpenRouter
Word Count
2,748
Company Posts That Month
24
Language
English
Hacker News Points
-
Post removed?
No
Summary

LLM-as-a-judge evaluation uses a separate language model to score an AI agent’s final responses against clear, written rubrics, addressing quality gaps that deterministic tests of tool calls or exact outputs may miss. It distinguishes candidate and judge models to reduce self-preference and other biases, and supports pointwise scoring for thresholds, pairwise comparisons for model or prompt selection, and reference-based checks for factual coverage. The guide recommends deterministic validators for exact, safety-critical requirements such as schemas, calculations, permissions, and tool side effects, while using judges for semantic qualities including grounding, completeness, instruction following, tone, and correct use of retrieved information. Using OpenRouter’s Ori Eval, developers can create TypeScript-based tests that combine observable assertions with rubric scoring, run evaluations in CI, compare candidate models, and control costs through sampling and focused criteria. Reliable judging requires observable rubrics, calibration against human-labeled examples, blinded evaluation inputs, version-controlled test data and configurations, awareness of output variance and known biases, and protection of sensitive production data.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 24 747 162 79 -85%
AI Agents 3 931 231 103 -84%
AI Guardrails 1 35 22 12 -94%
Harness engineering 1 33 23 14 -84%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.