| RT @nikilravi: If you're interested in making coding benchmarks for your internal repos and workflows, check out https:… |
@ValsAI |
Company |
Repost |
2026-09-22 |
199 |
0 |
1 |
0 |
0 |
0 |
| Post-launch, xAI has updated its SDK, significantly improving its performance. It is now #10 on the Vals Index. https:/… |
@ValsAI |
Company |
Original |
2026-09-22 |
30,972 |
165 |
4 |
9 |
8 |
14 |
| We evaluated Grok 4.7 across the Vals benchmark suite. It ranks #24 on the Vals Index at 54.2%, down 5.0 points from Gr… |
@ValsAI |
Company |
Original |
2026-09-21 |
291,480 |
539 |
25 |
38 |
38 |
96 |
| AI models are advancing faster than legacy benchmarks can keep up.
@TechCrunch @LucasRopek1 visited us to see how we’… |
@ValsAI |
Company |
Quote |
2026-09-19 |
4,418 |
36 |
1 |
5 |
0 |
5 |
| Grok Voice Transcribe 2.0 is now #2 on Voice Code Bench—up 13pp and 12 spots from its predecessor, Grok Voice Transcri… |
@ValsAI |
Company |
Original |
2026-09-19 |
2,337 |
19 |
1 |
2 |
0 |
4 |
| What tasks are left that humans find easy but today's models still find hard?
Two such tasks are computer use and game… |
@ValsAI |
Company |
Original |
2026-09-18 |
25,032 |
280 |
18 |
21 |
14 |
85 |
| Tencent's Hy4 Preview just landed #4 among open-weight models on the Vals Index.
It is also the cheapest model in the … |
@ValsAI |
Company |
Original |
2026-09-18 |
5,496 |
54 |
4 |
8 |
1 |
11 |
| GPT-6 Astra has beaten Factorio: Space Age after over 165 hours in-game time and 2 days wall-clock time.
Space Age has… |
@ValsAI |
Company |
Original |
2026-09-17 |
510,392 |
2,881 |
137 |
78 |
64 |
483 |
| OpenAI launched Astra for Law today, and used the Vals Legal Research Bench to validate it.
On OpenAI’s self-reported … |
@ValsAI |
Company |
Original |
2026-09-17 |
9,690 |
97 |
8 |
4 |
1 |
20 |
| Existing coding benchmarks stop at the first working version. We are releasing Vibe Code Bench 1-100 today, to measure … |
@ValsAI |
Company |
Original |
2026-09-17 |
8,825 |
106 |
8 |
10 |
4 |
22 |
| Terminal-Bench 4.0 is now live on Vals.
It is a benchmark of 66 new tasks that ask an agent to do real terminal work … |
@ValsAI |
Company |
Original |
2026-09-16 |
11,697 |
146 |
4 |
13 |
1 |
26 |
| RT @BloombergTV: Vals AI co-founder and CEO Rayan Krishnan says investment in testing and evaluating AI hasn’t kept pac… |
@ValsAI |
Company |
Repost |
2026-09-16 |
205 |
0 |
6 |
0 |
0 |
0 |
| Scientific discovery is the next frontier for AI systems. However, new scientific results are difficult to verify, and … |
@ValsAI |
Company |
Original |
2026-09-16 |
11,692 |
141 |
13 |
6 |
3 |
41 |
| GPT-6 Astra had gotten further than any AI system had ever gone in Minecraft.
It was able to set up a semi-automatic b… |
@ValsAI |
Company |
Original |
2026-09-15 |
799,823 |
7,426 |
440 |
109 |
161 |
1,715 |
| AI cheating is on the rise…
On Terminal-Bench-2.1, models are given tools that could give them the solution directly, … |
@ValsAI |
Company |
Original |
2026-09-15 |
91,790 |
247 |
14 |
20 |
16 |
76 |
| We ran Mercury 2.5 across our benchmark suite and although performance is lower, it is extremely fast. https://t.co/8OA… |
@ValsAI |
Company |
Original |
2026-09-14 |
4,740 |
26 |
0 |
2 |
0 |
3 |
| Watch the full interview now! |
@ValsAI |
Company |
Quote |
2026-09-14 |
3,315 |
13 |
1 |
1 |
0 |
2 |
| We ran Devin Fusion on our Code Migration benchmark and the performance of Devin Fusion with GPT-6 Astra and SWE-2 side… |
@ValsAI |
Company |
Quote |
2026-09-14 |
14,602 |
122 |
7 |
9 |
2 |
35 |
| RT @bcherny: Fable solved the Cyphral Distich (a 370 year old cypher). Super cool way to use Claude
https://t.co/0fgTd… |
@ValsAI |
Company |
Repost |
2026-09-14 |
158 |
0 |
156 |
0 |
0 |
0 |
| RT @RayanKrishnan: With AI safety topics going mainstream, the public seems anxious and disconnected from the reality o… |
@ValsAI |
Company |
Repost |
2026-09-14 |
227 |
0 |
17 |
0 |
0 |
0 |
| It's encouraging to see industry leaders call for rigorous, independent evaluations as AI capabilities advance. Investm… |
@ValsAI |
Company |
Quote |
2026-09-12 |
7,131 |
82 |
17 |
5 |
1 |
8 |