Decision model benchmark: Jev, Kev, Liquid d1, and more
Blog post from Arize
A benchmark of eight recently released decision models assessed their ability to detect hallucinations, evaluate LLM outputs, and route agent tool or skill choices, comparing accuracy, latency, cost, calibration, and deployment options across hosted APIs and locally run open models. Jev, Liquid d1, and Clef 27B were closely matched on core accuracy measures and approached the performance of a far more expensive LLM judge, while Jev generally offered strong speed and calibration among APIs, Kev emerged as the leading local model, and Clef 27B stood out as an open-weights option with consistently competitive accuracy. In the final routing test, Jev and Kev effectively tied, making the preferred deployment model—managed API or self-hosting—the main practical differentiator. A major finding was that several models performed very differently when equivalent evaluation questions were phrased with reversed polarity, with some producing confidently inverted results, underscoring that benchmarks and production evaluators should test prompt wording as carefully as model performance.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Jev | 40 | No monthly metrics for this publish month. | |||
| LLM | 9 | No monthly metrics for this publish month. | |||
| AI Guardrails | 2 | No monthly metrics for this publish month. | |||
| Observability | 2 | No monthly metrics for this publish month. | |||
| RAG | 1 | No monthly metrics for this publish month. | |||
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.