An Independent Evaluation of TypeSafe's Jev
Blog post from Vals
TypeSafe’s Jev is a non-generative model that returns probability distributions over predefined answer choices, designed for bounded classification tasks rather than open-ended writing. In tests against eleven LLMs, Jev matched several frontier systems on a 400-item SEC-filing claim-verification benchmark, scoring 0.975 accuracy at roughly $0.02 per 1,000 cases, while offering very low latency that changed little when many judgments were requested for the same document. However, it ranked last on a 396-question, 12-subtask LegalBench sample, particularly struggling with contract-entailment questions, illustrating that its strong verification performance did not generalize to all structured reasoning tasks. Its confidence calibration was best among tested systems on claim verification but worst on LegalBench, and a held-out routing experiment found it could automate 95% of verification cases at a 1.6% error rate under a nominal 1% target. The evaluation used fixed-answer datasets, human auditing, provider APIs, and paired comparisons, while noting limitations including the small LegalBench slice, possible benchmark contamination for LLMs, and uncertainty in selecting confidence thresholds.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Jev | 44 | No monthly metrics for this publish month. | |||
| GPT-6 Astra | 11 | No monthly metrics for this publish month. | |||
| LLM | 10 | No monthly metrics for this publish month. | |||
| Gemini 4 Argon | 9 | No monthly metrics for this publish month. | |||
| Opus 5.5 | 7 | No monthly metrics for this publish month. | |||
| GPT-6.1 Sol | 6 | No monthly metrics for this publish month. | |||
| Gemini 3.8 Flash | 2 | No monthly metrics for this publish month. | |||
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.