August 2026 Summaries
2 posts from Vals
Filter
Month:
Year:
Post Summaries
Back to Blog
Vals AI argues that AI development has advanced faster than independent evaluation methods, leaving existing benchmarks vulnerable to rapid saturation, training-data leakage, and conflicts of interest when model developers assess their own systems. The company aims to provide independent, reproducible benchmarks for professional tasks in fields such as law, finance, engineering, and medicine, using private test sets and partnerships with domain institutions to maintain measurement integrity. It says its evaluations have been cited by major AI labs, used by enterprises selecting models, and supported AI policy work in government, while reporting significant revenue, customer, and team growth. Vals announced a $40 million Series A funding round at a $400 million valuation led by Andreessen Horowitz, alongside the general availability of Vals Smith for creating coding benchmarks from GitHub repositories, new frontier-risk benchmarks including cyber and mental-health work, and a rebuilt Vals platform and expanded index.
Aug 13, 2026
524 words in the original blog post.
Vals AI has launched an initiative to develop independent evaluations of how AI systems handle mental-health-related interactions with children and teenagers, citing rising youth use of chatbots and companion AI alongside concerns about potential harms. The effort follows research indicating that many young people seek mental health advice or companionship from AI, legal cases alleging chatbot-related harm to minors, and studies showing that models can fail in longer or nuanced conversations despite performing well on direct safety prompts. Vals is working with clinicians and academic researchers to assess self-harm and crisis responses, inappropriate assumptions of therapeutic or authoritative roles, and unhealthy emotional dependence, with an emphasis on how risks develop across multi-turn exchanges. Early findings suggest that models generally recognize explicit suicide disclosures but vary significantly in the usefulness of their follow-up, while more subtle risks can emerge through ordinary questions about health, treatment, or support and through gradually developing relational dynamics. The initiative comes as AI companies introduce youth-focused safeguards and policymakers pursue reporting, risk-assessment, and independent auditing requirements, and Vals plans to share detailed findings with policymakers before publishing further methodology and results.
Aug 12, 2026
715 words in the original blog post.