Home / Companies / Harness / Blog / Post Details
Content Deep Dive

Catch AI Regressions Before They Ship with AI Evals in CI/CD | Harness Blog

Blog post from Harness

Post Details
Company
Date Published
Author
Shibam Dhar
Word Count
969
Company Posts That Month
2
Language
English
Hacker News Points
-
Post removed?
No
Summary

Harness AI Evals is presented as a CI/CD quality gate for AI agents, addressing behavioral failures that conventional software tests may miss even when services are available and functioning technically. In an e-commerce support-agent example, 32 golden scenarios were evaluated for answer relevancy, task completion, and toxicity, with deployments blocked unless they met a 70% threshold. An initial 65% result exposed incorrect billing facts, incomplete shipping information, indirect policy answers, and responses to the wrong customer question, leading developers to improve knowledge sources and prompts rather than lower the standard. Subsequent runs achieved 75% and 78%, but differing results across identical evaluations also highlighted AI non-determinism and the need to verify improvements through repeated testing while distinguishing genuine quality failures from broken infrastructure or evaluators. By placing evaluation directly within the delivery pipeline, Harness aims to make AI behavior a release criterion alongside builds, tests, and service health checks.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Agents 3 No monthly metrics for this publish month.
Developer Experience 1 No monthly metrics for this publish month.
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.