Home / Companies / Vals / Blog / Post Details
Content Deep Dive

An Independent Evaluation of TypeSafe's Jev

Blog post from Vals

Post Details
Company
Date Published
Author
Connor Frank
Word Count
2,766
Company Posts That Month
7
Language
English
Hacker News Points
1
Post removed?
No
Summary

TypeSafe’s Jev is a non-generative model that returns probability distributions over predefined answer choices, designed for bounded classification tasks rather than open-ended writing. In tests against eleven LLMs, Jev matched several frontier systems on a 400-item SEC-filing claim-verification benchmark, scoring 0.975 accuracy at roughly $0.02 per 1,000 cases, while offering very low latency that changed little when many judgments were requested for the same document. However, it ranked last on a 396-question, 12-subtask LegalBench sample, particularly struggling with contract-entailment questions, illustrating that its strong verification performance did not generalize to all structured reasoning tasks. Its confidence calibration was best among tested systems on claim verification but worst on LegalBench, and a held-out routing experiment found it could automate 95% of verification cases at a 1.6% error rate under a nominal 1% target. The evaluation used fixed-answer datasets, human auditing, provider APIs, and paired comparisons, while noting limitations including the small LegalBench slice, possible benchmark contamination for LLMs, and uncertainty in selecting confidence thresholds.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Jev 44 No monthly metrics for this publish month.
GPT-6 Astra 11 No monthly metrics for this publish month.
LLM 10 No monthly metrics for this publish month.
Gemini 4 Argon 9 No monthly metrics for this publish month.
Opus 5.5 7 No monthly metrics for this publish month.
GPT-6.1 Sol 6 No monthly metrics for this publish month.
Gemini 3.8 Flash 2 No monthly metrics for this publish month.
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.