Home / Companies / Confident AI / Blog / Post Details
Content Deep Dive

How I Built Deterministic LLM Evaluation Metrics for DeepEval

Blog post from Confident AI

Post Details
Company
Date Published
Author
Jeffrey Ip
Word Count
2,335
Company Posts That Month
1
Language
English
Hacker News Points
-
Post removed?
No
Summary

The author of the text noticed a divide among DeepEval users who were either happy with the out-of-the-box metrics or not. The issue was that the built-in metrics didn't fit their use case and weren't deterministic enough, leading to hundreds of lines of code dedicated to tweaking evaluation logic. To address this, the author introduced a new metric called DAG (Deep Acyclic Graph) which is structured around LLM-powered decision trees, providing customizability and determinism for evaluations. The DAG metric breaks down an LLM test case into atomic units and uses four core node types: Task nodes, Binary Judgment nodes, Non-Binary Judgment nodes, and Verdict nodes. This allows users to easily build DAGs within DeepEval, making evaluation easy for even smaller models to handle, and benefits from optimized parallel execution, efficient cost management, built-in caching, and error handling. The author concludes that the DAG metric solves the core problem of traditional metrics lacking control and provides a transparent, efficient, and adaptable evaluation process.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 42 4,013 569 191 -13%
AI Guardrails 8 242 83 45 -30%
RAG 2 1,528 261 92 -30%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.