Home / Companies / Vals / Blog / Post Details
Content Deep Dive

Measuring Healthcare AI Where It Actually Works

Blog post from Vals

Post Details
Company
Date Published
Author
Rayan Krishnan
Word Count
559
Company Posts That Month
1
Language
English
Hacker News Points
-
Post removed?
No
Summary

MedScribe and MedCode are new healthcare benchmarks designed to assess AI performance on real clinical workflows using de-identified patient records, addressing a gap in evaluations focused mainly on general medical knowledge. MedScribe tests the generation of SOAP notes from physician-patient conversations and found that leading models achieved 85–88% accuracy, with GPT 5.1 scoring 88.09%, suggesting AI can meaningfully support documentation tasks that contribute substantially to clinician workload and burnout. MedCode evaluates ICD-10-CM diagnosis coding across 2,755 samples and produced much weaker results, with leading models generally scoring 49–53% and Gemini 3 Flash reaching 55.92%. The 32-percentage-point difference reflects the contrast between documentation as a summarization and organization task and coding as a precise, hierarchical, rule-driven process with reimbursement, compliance, and audit consequences. Models performed relatively better on common physical conditions than on mental health diagnoses, highlighting the need for further development before coding can be reliably automated.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.