Home / Companies / Vals / Blog / Post Details
Content Deep Dive

Has the Bitter Lesson Come for AI Detectors?

Blog post from Vals

Post Details
Company
Date Published
Author
Daniel Fein
Word Count
611
Company Posts That Month
6
Language
English
Hacker News Points
-
Post removed?
No
Summary

Vals’ AI Detection Benchmark evaluates specialized AI-writing detectors and general-purpose language models on a private dataset of human-authored documents created before November 2022 and AI-assisted rewrites that altered more than half of each original passage. Claude Opus 5.5 achieved the highest reported balanced accuracy at 98.38%, followed by GPT-6 Astra at 95.75%, with Astra costing less per document, while results showed that identifying rewritten AI text was generally more difficult than recognizing human writing. The benchmark also highlights tradeoffs among detectors, as GPTZero reportedly classified all human documents correctly but detected only 56.3% of AI rewrites, whereas Sapling detected more rewrites but misclassified more human work. Adversarial testing with rewrites generated by frontier models found that such edits can evade reliable detection across systems, although the generating models performed better against their own outputs than Pangram. The report cautions that its limited, deliberately controlled dataset and adversarial prompts may not represent real-world AI-detection performance across broader writing styles, human-AI collaboration patterns, and model distributions.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
GPT-6 Astra 2 No monthly metrics for this publish month.
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.