Home / Companies / Vals / Blog / October 2026

October 2026 Summaries

3 posts from Vals

Filter
Month: Year:
Post Summaries Back to Blog
TypeSafe’s Jev is a non-generative model that returns probability distributions over predefined answer choices, designed for bounded classification tasks rather than open-ended writing. In tests against eleven LLMs, Jev matched several frontier systems on a 400-item SEC-filing claim-verification benchmark, scoring 0.975 accuracy at roughly $0.02 per 1,000 cases, while offering very low latency that changed little when many judgments were requested for the same document. However, it ranked last on a 396-question, 12-subtask LegalBench sample, particularly struggling with contract-entailment questions, illustrating that its strong verification performance did not generalize to all structured reasoning tasks. Its confidence calibration was best among tested systems on claim verification but worst on LegalBench, and a held-out routing experiment found it could automate 95% of verification cases at a 1.6% error rate under a nominal 1% target. The evaluation used fixed-answer datasets, human auditing, provider APIs, and paired comparisons, while noting limitations including the small LegalBench slice, possible benchmark contamination for LLMs, and uncertainty in selecting confidence thresholds.
Oct 06, 2026 2,766 words in the original blog post.
Luttinger-compensated magnets are proposed as a promising middle ground between ferromagnets, which provide spin-polarized electrons but produce disruptive stray fields and switch relatively slowly, and conventional antiferromagnets, which have no net magnetic field and can switch quickly but lack useful spin sorting. These materials retain zero net magnetic moment because opposing atomic spins cancel, while inequivalent atomic sites allow electron states near their semiconductor band edges to remain spin-polarized, potentially enabling dense, fast spintronic memory devices. Quantum-mechanical simulations identified two candidates: newly designed YBaMnFeO₅, predicted to have a 2.35 eV band gap, large spin-sorted energy windows, and magnetic order above room temperature, but likely difficult to synthesize because its required manganese–iron ordering may break down at high processing temperatures; and KV[Cr(CN)₆], a Prussian-blue-related compound first made in 1999, predicted to have a roughly 2.1 eV band gap and robust spin sorting while experimentally retaining magnetic order up to 376 K. Although the latter’s spin polarization and ideal zero moment still require direct measurement, its chemically locked crystal structure and room-temperature magnetic behavior make it a particularly notable candidate for practical spin-based technologies, with computational data and methods publicly shared for verification.
Oct 04, 2026 1,771 words in the original blog post.
Evaluating AI web-search tools is difficult because static benchmarks can be contaminated, may reward models’ pretrained knowledge rather than search use, and often contain narrow tasks unlike real professional work. The passage argues that widely used benchmarks such as BrowseComp and HLE have become less reliable as questions and answers leak online, while research suggests some models can answer many BrowseComp questions without tools and may use search chiefly to verify existing knowledge. It describes alternative efforts by search providers to use newer or dynamically generated tasks, then introduces the Vals Web Search Index as an independent evaluation intended to measure whether an agent can complete realistic knowledge-work tasks using a specific search tool. The index holds the model and evaluation harness constant while swapping search tools, scores final end-to-end answers rather than individual search-result rankings, and currently draws from expert-authored finance and legal benchmarks. Its private, continuously refreshed test set and zero-data-retention procedures are intended to limit memorization and leakage; reported no-search results were much lower than search-enabled performance. The authors contend that benchmarks should prioritize accurate completion of consequential real-world tasks over retrieval of isolated facts, since evaluation standards influence how search systems are optimized.
Oct 02, 2026 1,712 words in the original blog post.