Home / Companies / Arize / Blog / Post Details
Content Deep Dive

Decision model benchmark: Jev, Kev, Liquid d1, and more

Blog post from Arize

Post Details
Company
Date Published
Author
Jim Bennett
Word Count
2,802
Company Posts That Month
9
Language
English
Hacker News Points
-
Post removed?
No
Summary

A benchmark of eight recently released decision models assessed their ability to detect hallucinations, evaluate LLM outputs, and route agent tool or skill choices, comparing accuracy, latency, cost, calibration, and deployment options across hosted APIs and locally run open models. Jev, Liquid d1, and Clef 27B were closely matched on core accuracy measures and approached the performance of a far more expensive LLM judge, while Jev generally offered strong speed and calibration among APIs, Kev emerged as the leading local model, and Clef 27B stood out as an open-weights option with consistently competitive accuracy. In the final routing test, Jev and Kev effectively tied, making the preferred deployment model—managed API or self-hosting—the main practical differentiator. A major finding was that several models performed very differently when equivalent evaluation questions were phrased with reversed polarity, with some producing confidently inverted results, underscoring that benchmarks and production evaluators should test prompt wording as carefully as model performance.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Jev 40 No monthly metrics for this publish month.
LLM 9 No monthly metrics for this publish month.
AI Guardrails 2 No monthly metrics for this publish month.
Observability 2 No monthly metrics for this publish month.
RAG 1 No monthly metrics for this publish month.
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.