Is Jev as Accurate as Frontier Models at Classification?
Blog post from OpenRouter
A comparison of TypeSafe’s Jev 1.13 decision model and Claude Opus 5 on the 3,080-example Banking77 intent-classification test found that Jev achieved 81.0% accuracy and 80.5% macro-F1, trailing Opus’s 84.4% accuracy and 83.6% macro-F1 by 3.3 percentage points. Both models classified messages into 77 banking-support intents using identical label criteria, with no malformed outputs, but Jev delivered substantially lower latency and cost, recording a 175 ms median response time and $0.11 per thousand requests compared with Opus’s 2,266 ms and $2.42 per thousand when prompt caching was used. Jev’s confidence scores were not fully calibrated but effectively ranked likely errors, enabling a cascade in which requests below 0.90 confidence were routed to Opus; this preserved 84.0% accuracy, within 0.4 points of Opus alone, while reducing costs to $0.69 per thousand requests. The findings are limited to one dataset, domain, prompt design, and short test period, and both models underperformed fine-tuned encoders, partly because some intent labels were ambiguous when described only by name.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.