Home / Companies / TestMu AI / Blog / Post Details
Content Deep Dive

Evaluating LLM Relevancy with DeepEval [Testμ 2026]

Blog post from TestMu AI

Post Details
Company
Date Published
Author
TestMu AI
Word Count
2,414
Company Posts That Month
54
Language
English
Hacker News Points
-
Post removed?
No
Summary

At Testμ Conf 2026, Salesforce engineer Monika Sharma explained how DeepEval, an open-source Python framework often described as “pytest for LLMs,” evaluates non-deterministic LLM responses through semantic scoring by a judge model rather than exact string matching. DeepEval provides RAG metrics including faithfulness, answer relevancy, contextual relevancy, precision, and recall, along with agent, safety, conversational, and customizable G-Eval metrics, producing scores from zero to one, pass/fail results, and written rationales. Sharma recommended starting thresholds around 0.5 and raising them as model behavior becomes stable, while noting that a “none” result generally signals an output-format parsing problem rather than poor quality. She positioned evals as either unit or integration tests depending on whether they assess direct agent APIs or UI-based experiences, and said they can be incorporated into standard CI/CD pipelines. However, she emphasized that human judgment remains necessary for multimodal outputs such as chatbot responses containing images, emoticons, and feedback controls, which text-focused evaluation may not understand. Production autonomy should be supported by many varied utterances and consistent results across repeated runs, not isolated passing tests, while agents should generally have read-only access to knowledge bases for security reasons.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 11 747 162 79 -85%
RAG 5 101 30 23 -91%
AI Agents 1 931 231 103 -84%
Subagents 1 15 10 7 -95%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.