| We let open models access the outside world during evals on tasks from public benchmarks. Here's what happened:
Most w… |
@AI21Labs |
Company |
Original |
2026-09-29 |
14,184 |
7 |
3 |
7 |
1 |
2 |
| Catch up with our two talks from @aiDotEngineer World's Fair, now available online👇
Stop chunking like it's 2022: http… |
@AI21Labs |
Company |
Quote |
2026-09-28 |
12,985 |
4 |
0 |
1 |
1 |
0 |
| RT @YuvalinTheDeep: My talk from @aiDotEngineer SF is up!
While everyone moved on to agents, somebody still has to tak… |
@AI21Labs |
Company |
Repost |
2026-09-24 |
114 |
0 |
1 |
0 |
0 |
0 |
| RT @alexwg: Conversations in Action episode #14 with @Stanford emeritus professor and @AI21Labs co-founder Yoav Shoham … |
@AI21Labs |
Company |
Repost |
2026-09-22 |
110 |
0 |
9 |
0 |
0 |
0 |
| Sometimes all your agents are cooking at once. So you build Claudino 🦖
A one-key terminal game to play while Claude … |
@AI21Labs |
Company |
Original |
2026-09-08 |
1,076 |
9 |
0 |
2 |
1 |
0 |
| Independent verifiers improve agent output - but frontier verifiers are expensive. So we trained our own.
Results on @… |
@AI21Labs |
Company |
Original |
2026-09-02 |
787 |
10 |
3 |
3 |
0 |
2 |
| Adding a verifier improves every agentic AI system. But frontier verifiers are expensive - so we trained our own. Match… |
@AI21Labs |
Company |
Quote |
2026-08-23 |
17,882 |
20 |
5 |
8 |
1 |
8 |
| RT @YuvalinTheDeep: 1/ Your agent already found the right answer, but your aggregation threw it away. Across every gene… |
@AI21Labs |
Company |
Repost |
2026-08-20 |
101 |
0 |
2 |
0 |
0 |
0 |
| AI21 joined @nvidia, @Microsoft , @a16z, and dozens of others in signing the Open Weights and American AI Leadership le… |
@AI21Labs |
Company |
Quote |
2026-07-28 |
2,826 |
13 |
3 |
0 |
0 |
2 |
| New episode drop: Thrilled to welcome Not Diamond co-founder & CEO @tomas_hk to YAAP, hosted by @YuvalinTheDeep, to… |
@AI21Labs |
Company |
Original |
2026-07-21 |
1,915 |
15 |
4 |
5 |
1 |
1 |
| 1/3 New SOTA on the full 731-task SWE-Bench Pro: 80.8% resolve at $5.99/task by building our coding agent pipeline like… |
@AI21Labs |
Company |
Original |
2026-07-16 |
1,930 |
24 |
1 |
5 |
1 |
4 |
| Your agents are made of models + harnesses. So is your bill.
@Get_Writer's "Harness Effect" study held models fixed, … |
@AI21Labs |
Company |
Original |
2026-07-15 |
630 |
4 |
2 |
1 |
0 |
2 |
| 1/3 Best-of-N leaves $$ on the table by not accounting for variance in task difficulty. We built budget-aware execution… |
@AI21Labs |
Company |
Original |
2026-07-08 |
890 |
8 |
2 |
1 |
0 |
1 |
| 1/3 We just landed #1 on DeepResearch Bench II without building a single new agent. https://t.co/fvuBGFM3HN |
@AI21Labs |
Company |
Original |
2026-06-24 |
1,377 |
9 |
2 |
2 |
0 |
1 |
| 1/5 Our latest Labs in Front piece: Agent pipeline order matters. By reversing a common agent recipe - scale first, enr… |
@AI21Labs |
Company |
Original |
2026-06-04 |
1,741 |
8 |
2 |
5 |
0 |
0 |
| 1/5 We’re seeing 4 common agent optimization methods for hitting the right accuracy-cost or accuracy-latency tradeoff. … |
@AI21Labs |
Company |
Original |
2026-05-13 |
986 |
6 |
1 |
3 |
0 |
2 |
| New #YAAP episode out now 🎙️
@yuvalinthedeep sits down with @mikegchambers from @awsdevelopers to unpack harness engin… |
@AI21Labs |
Company |
Original |
2026-05-07 |
1,061 |
6 |
0 |
1 |
0 |
0 |
| Join @YuvalinTheDeep, Senior Developer Advocate at @AI21Labs, for a live webinar in partnership with @DataCamp: The Fou… |
@AI21Labs |
Company |
Original |
2026-05-05 |
1,015 |
3 |
1 |
0 |
1 |
0 |
| 1/5 We hit SOTA performance on BrowseComp-Plus with 95.18% accuracy using AI21 Maestro’s agent optimization.
Here’s h… |
@AI21Labs |
Company |
Original |
2026-04-30 |
1,770 |
8 |
2 |
2 |
0 |
3 |
| Live from @DeepLearningAI conference: our CPO Or Dagan is taking the stage, explaining how we got SOTA on Browsecomp-Pl… |
@AI21Labs |
Company |
Original |
2026-04-29 |
119,093 |
23 |
2 |
1 |
0 |
1 |
| Day 2 at AI Dev 26 in SF surrounded by the best builders.
📍Find us at booth 121.
@DeepLearningAI https://t.co/yd9wCN… |
@AI21Labs |
Company |
Original |
2026-04-29 |
1,066 |
3 |
0 |
0 |
0 |
0 |
| Day 1 of #AIDevSF is live. 🙌
📍 Find us at booth 121 🎤 Catch our Chief Product Officer Or Dagan's session on Day 2: Eff… |
@AI21Labs |
Company |
Original |
2026-04-28 |
1,352 |
9 |
3 |
0 |
0 |
0 |
| Tune in to Or Dagan on the Life Self Mastery podcast with @rohitmal, unpacking why most enterprise AI projects fail bef… |
@AI21Labs |
Company |
Original |
2026-04-27 |
1,625 |
2 |
0 |
1 |
1 |
0 |
| Are you coming to @DeepLearningAI AI's AI Dev 26 x SF next week?
Our Chief Product & Strategy Officer, Or Dagan, will… |
@AI21Labs |
Company |
Original |
2026-04-23 |
84,884 |
7 |
1 |
0 |
1 |
0 |
| We just hit #1 on the @huggingface BrowseComp-Plus leaderboard.
Best accuracy: 92.53%. Best recall: 88.79%. Lowest cal… |
@AI21Labs |
Company |
Original |
2026-04-22 |
1,810 |
8 |
1 |
0 |
0 |
0 |
| Attending #AIDev26 by @DeepLearningAI?
Join @AI21Labs, @trychroma + @Baseten for a panel on optimizing modern AI syste… |
@AI21Labs |
Company |
Original |
2026-04-20 |
1,480 |
5 |
0 |
0 |
0 |
0 |
| 1/5 When we saw our Reducer (=LLM judge component in Maestro, our agentic framework, that selects the best output from … |
@AI21Labs |
Company |
Original |
2026-04-15 |
1,812 |
11 |
2 |
1 |
1 |
4 |
| Catch Or Dagan, AI21's Chief Product & Strategy Officer, at #AIDev26 on April 29: "An End to Manual Tinkering: Opti… |
@AI21Labs |
Company |
Original |
2026-04-14 |
1,187 |
5 |
0 |
0 |
1 |
0 |
| AI Dev 26 brings the builders together. We'll be among them.
Meet us there 👇 @DeepLearningAI @AndrewYNg |
@AI21Labs |
Company |
Quote |
2026-04-12 |
8,044 |
15 |
6 |
2 |
0 |
3 |
| Routing every task to your largest model burns tokens, adds latency, and inflates costs.
@AI21Labs' Maestro Orchestra… |
@AI21Labs |
Company |
Original |
2026-04-09 |
1,116 |
8 |
1 |
0 |
0 |
2 |
| Most agents guess the next move. They don't search.
Production systems need to branch, route, and kill bad paths befor… |
@AI21Labs |
Company |
Original |
2026-04-07 |
1,235 |
5 |
1 |
0 |
1 |
3 |
| You're overpaying for inference. SWE-bench shows cheaper models solve the same easy problems as frontier ones.
The iss… |
@AI21Labs |
Company |
Original |
2026-04-06 |
1,126 |
3 |
2 |
0 |
1 |
0 |
| Congratulations to our Co-founder and Co-CEO @yshoham, named an @aaas Fellow for pioneering contributions to agentic AI… |
@AI21Labs |
Company |
Original |
2026-03-30 |
1,348 |
3 |
0 |
1 |
1 |
0 |
| 1/4 We hit a strange logprob mismatch while training Jamba 3B with GRPO.
Rollout logprobs and training-side recompute… |
@AI21Labs |
Company |
Original |
2026-03-26 |
1,160 |
6 |
1 |
1 |
1 |
1 |
| Most AI demos work. Most AI systems don't.
SEAL (@AI21Labs' Solutions Engineering & Architecture Lab) exists to cl… |
@AI21Labs |
Company |
Original |
2026-03-24 |
852 |
4 |
1 |
2 |
0 |
0 |
| Day 2 at @nvidia GTC 26. Stop by booth 3103 to meet the team and learn more about what we’re building for enterprise AI… |
@AI21Labs |
Company |
Original |
2026-03-17 |
837 |
8 |
0 |
1 |
0 |
0 |
| We are heading to @NVIDIA #GTC26 in San Jose.
If you’re attending, stop by Booth 3103 to meet the team and see what we… |
@AI21Labs |
Company |
Original |
2026-03-12 |
666 |
3 |
0 |
0 |
0 |
0 |
| Standard RAG falls apart on aggregative queries.
Things like "average ARR for companies with >1k employees" require r… |
@AI21Labs |
Company |
Original |
2026-03-10 |
867 |
7 |
1 |
0 |
0 |
5 |
| Test-time compute scaling for agents isn’t just running multiple full trajectories and picking a winner. Duplicate a 50… |
@AI21Labs |
Company |
Original |
2026-03-05 |
958 |
3 |
0 |
0 |
0 |
0 |
| 1/4 Our Head of Human Alignment and Standards @JulieFadlon, PhD, breaks down what agent architectures have to learn fro… |
@AI21Labs |
Company |
Original |
2026-02-26 |
938 |
7 |
2 |
1 |
0 |
1 |
| 1/4 Our Head of Human Alignment and Standards @JulieFadlon, PhD, breaks down what agent architectures have to learn fro… |
@AI21Labs |
Company |
Original |
2026-02-26 |
361 |
3 |
0 |
0 |
0 |
1 |
| https://t.co/FJu144rnxr |
@AI21Labs |
Company |
Original |
2026-02-24 |
1,040 |
5 |
0 |
1 |
0 |
0 |
| Dear AI agent,
Please don’t get creative with our Q4 reports. We might go to jail.
Your Finance team.
👉 Build Boring … |
@AI21Labs |
Company |
Original |
2026-02-16 |
778 |
4 |
0 |
0 |
1 |
0 |
| 1/5 As part of our work on improving the efficiency of our LLM online-RL training pipelines, we cut policy update step … |
@AI21Labs |
Company |
Original |
2026-02-11 |
680 |
6 |
1 |
1 |
0 |
2 |
| 1/5 Go Big or Go OOM: The Art of Scaling vLLM 🎯.
We doubled throughput and cut latency in half-same GPUs, just better v… |
@AI21Labs |
Company |
Original |
2026-02-09 |
16,342 |
41 |
4 |
2 |
1 |
30 |
| @AnthropicAI: “The longer models spend reasoning and taking actions, the more incoherent they become.”
Our boring ag… |
@AI21Labs |
Company |
Original |
2026-02-08 |
370 |
3 |
0 |
1 |
0 |
0 |
| @AnthropicAI: “The longer models spend reasoning and taking actions, the more incoherent they become.”
Our boring agen… |
@AI21Labs |
Company |
Original |
2026-02-05 |
320 |
9 |
1 |
0 |
0 |
1 |
| We asked our boring agents about Moltbook, the new social network for agents. They did not approve. 🦞 🚫
#Boringagents … |
@AI21Labs |
Company |
Original |
2026-02-02 |
1,254 |
15 |
2 |
3 |
0 |
2 |
| 1/5 Debugging vLLM: The silent corruption bug
1/1000 Jamba generations collapsed into confident gibberish during RL tra… |
@AI21Labs |
Company |
Original |
2026-01-29 |
984 |
7 |
0 |
1 |
0 |
3 |
| This is Linda.
25,000 regulation docs. Before lunch. No smile.
She’s a custom AI agent for banking compliance.
Consis… |
@AI21Labs |
Company |
Original |
2026-01-28 |
727 |
9 |
2 |
0 |
0 |
0 |
| Boring AI Agents? They aren't players, they are workflow slayers:
- Oliver: Humor in beta. Grounded in data.
- Nancy: … |
@AI21Labs |
Company |
Original |
2026-01-20 |
2,488 |
20 |
2 |
5 |
2 |
3 |
| “Parallel agents are easy when they’re read-only. It gets complicated the moment they change files, send emails, or tou… |
@AI21Labs |
Company |
Original |
2026-01-20 |
606 |
7 |
1 |
1 |
0 |
1 |
| “If your AI is wrong 2% of the time, you’re not in the game.”
In his interview with @vladdoes, our Co-Founder & Co-CEO… |
@AI21Labs |
Company |
Original |
2026-01-19 |
1,104 |
13 |
3 |
0 |
1 |
4 |
| New benchmark signal: Jamba2 places high on @vectara's HHEM leaderboard for hallucination rate competitive with open mo… |
@AI21Labs |
Company |
Original |
2026-01-14 |
816 |
10 |
1 |
1 |
0 |
0 |
| Meet our boring AI agents.
They don’t invent facts.
They’re dull in chats. Cold. Distant. Consistent.
Built for ent… |
@AI21Labs |
Company |
Original |
2026-01-13 |
195,664 |
91 |
6 |
6 |
2 |
51 |
| 1/5 We ran SWE-bench 200,000+ times to get statistical confidence in agentic evals.
The main lesson wasn’t about prom… |
@AI21Labs |
Company |
Original |
2026-01-12 |
1,785 |
6 |
0 |
1 |
1 |
1 |
| 1/4 MCP works great… until you run multiple subagents on the same task and they need to write files (not just read).
P… |
@AI21Labs |
Company |
Original |
2026-01-09 |
857 |
6 |
1 |
1 |
1 |
3 |
| 1/4 🚀Introducing Jamba2, a memory-efficient open source model family built for total enterprise reliability and steerab… |
@AI21Labs |
Company |
Original |
2026-01-08 |
3,303 |
25 |
9 |
1 |
4 |
2 |
| 1/6 Long-horizon agentic tasks are breaking our mental models. More tokens, bigger models, and best-of-N only go so far… |
@AI21Labs |
Company |
Original |
2026-01-07 |
1,554 |
18 |
6 |
1 |
2 |
3 |
| 🎙️ How do you make a high quality deep research agent without burning thousands of tokens?
@tavilyai's Dean Sacoransky… |
@AI21Labs |
Company |
Original |
2026-01-06 |
489 |
6 |
0 |
0 |
0 |
0 |
| 🎙️ In this episode of Yet Another AI Podcast, we sat down with @lindavivah from @anyscalecompute on scaling AI workload… |
@AI21Labs |
Company |
Original |
2025-12-30 |
2,684 |
13 |
2 |
0 |
0 |
10 |
| Enterprises moved fast on AI in 2025. Only a few figured out how to scale it. The data explains why. https://t.co/MUwzM… |
@AI21Labs |
Company |
Original |
2025-12-29 |
765 |
5 |
2 |
1 |
0 |
1 |
| Happy holidays from @AI21Labs. Thanks to our customers, partners, and team for helping to make this year one for the bo… |
@AI21Labs |
Company |
Original |
2025-12-23 |
621 |
2 |
0 |
0 |
0 |
0 |
| 🎙️New episode of Yet Another AI Podcast
@AiImagen shares how to build real moats when everyone uses the same models an… |
@AI21Labs |
Company |
Original |
2025-12-17 |
1,045 |
7 |
0 |
0 |
0 |
0 |
| Vibe Agent in AI21 Maestro helps you create AI agents from a single plain-English description. It suggests purpose, val… |
@AI21Labs |
Company |
Original |
2025-12-16 |
926 |
4 |
0 |
0 |
0 |
1 |
| @GroupeFnacDarty selects @AI21Labs as a strategic AI partner to tackle retail’s most complex challenges, starting with … |
@AI21Labs |
Company |
Original |
2025-12-10 |
748 |
2 |
0 |
0 |
0 |
1 |
| AI21 Maestro is @AI21Labs' orchestration platform for building reliable end-to-end AI workflows.
It combines multi-ste… |
@AI21Labs |
Company |
Original |
2025-12-08 |
999 |
14 |
1 |
0 |
0 |
1 |
| 🚀 Big news for the industry: AI21 Maestro is now available for deployment in your @awscloud VPC , seamlessly integratin… |
@AI21Labs |
Company |
Original |
2025-12-04 |
918 |
4 |
1 |
0 |
0 |
1 |
| We took to the floor at @awscloud re:Invent to ask when people last cursed at AI.
Let’s just say… the answers did not … |
@AI21Labs |
Company |
Original |
2025-12-03 |
1,016 |
3 |
0 |
1 |
0 |
0 |
| Live from @awscloud re:Invent in Las Vegas. We are on-site at Booth #1213 and ready to show how accurate, reliable know… |
@AI21Labs |
Company |
Original |
2025-12-02 |
1,273 |
5 |
1 |
0 |
0 |
0 |
| We’re taking a major step toward making AI more open, reliable & enterprise-ready.
@AI21Labs is partnering with @togeth… |
@AI21Labs |
Company |
Original |
2025-11-26 |
547 |
5 |
0 |
0 |
0 |
0 |
| Heading to @awscloud #reInvent? Join @AI21Labs & @deepchecks for an AI agents meetup on Dec 4 in Las Vegas. Learn f… |
@AI21Labs |
Company |
Original |
2025-11-25 |
605 |
4 |
0 |
1 |
0 |
1 |
| We have been recognized as an *Emerging Visionary* in two @Gartner_inc® Emerging Market Quadrants: Generative AI Engine… |
@AI21Labs |
Company |
Original |
2025-11-24 |
1,346 |
3 |
1 |
2 |
1 |
0 |
| “AI21 Maestro’s accuracy fix for RAG’s blind spots”.
Awesome to see @PaulBaier and the @GAIinsights team highlight ou… |
@AI21Labs |
Company |
Original |
2025-11-17 |
852 |
9 |
0 |
1 |
0 |
3 |
| 📡 Live from AI Dev 25 x NYC!
Our Co-founder & CEO @origoshen is on stage right now discussing why reliability is th… |
@AI21Labs |
Company |
Original |
2025-11-14 |
733 |
4 |
1 |
0 |
0 |
0 |
| 🚀 Introducing Structured RAG (S-RAG)
S-RAG transforms unstructured data into a structured, query-aware representation.… |
@AI21Labs |
Company |
Original |
2025-11-12 |
35,830 |
232 |
33 |
6 |
3 |
214 |
| New episode of Yet Another AI Podcast 🎙
Yuval talks with Henry Yin (Co-founder & CTO of @agihouse_org) about “The Hous… |
@AI21Labs |
Company |
Original |
2025-11-11 |
568 |
4 |
0 |
0 |
0 |
1 |
| How much RAM do you need to run tiny models?
Jamba Reasoning 3B runs on just 2.25 GiB, the lightest among small models… |
@AI21Labs |
Company |
Original |
2025-11-06 |
607 |
6 |
0 |
2 |
0 |
4 |
| Our Jamba Reasoning 3B model took on Qwen 3 4B 2507 by @Alibaba_Cloud in the same 60K token task. Jamba Reasoning 3B fi… |
@AI21Labs |
Company |
Original |
2025-11-04 |
520 |
8 |
0 |
0 |
0 |
1 |
| We wrapped an incredible day yesterday at @githubuniverse ’25. Day 2 - here we go!
#GitHubUniverse @GitHubCommunity htt… |
@AI21Labs |
Company |
Original |
2025-10-29 |
472 |
5 |
0 |
1 |
0 |
1 |
| 🕸️ The web is the world’s biggest database but scraping its data at scale is tricky. On this week’s #YAAP episode, @Yuv… |
@AI21Labs |
Company |
Original |
2025-10-28 |
710 |
6 |
1 |
0 |
0 |
1 |
| A week with the @nvidia DGX Spark and we’re already running private agents trained on private data with Jamba 3B Reason… |
@AI21Labs |
Company |
Original |
2025-10-27 |
1,040 |
14 |
2 |
2 |
0 |
0 |
| We’re joining @jfrog's MLOps Days on Oct 22 in Tel Aviv!
Catch @CohenShuki, VP of Data at @AI21Labs, sharing how small… |
@AI21Labs |
Company |
Original |
2025-10-16 |
784 |
2 |
1 |
1 |
0 |
0 |
| 🚀 @VentureBeat spotlights Jamba Reasoning 3B, our tiny open-source model that outperforms others across reasoning bench… |
@AI21Labs |
Company |
Original |
2025-10-10 |
720 |
5 |
0 |
0 |
0 |
1 |
| 🧠 Jamba Reasoning 3B leads tiny reasoning models (Artificial Analysis).
🥇 #1 on #IFBench (52%) for instruction followin… |
@AI21Labs |
Company |
Original |
2025-10-09 |
1,257 |
16 |
2 |
1 |
0 |
0 |
| 1/5 Releasing Jamba Reasoning 3B under Apache 2.0: Hybrid SSM-Transformer architecture that tops accuracy & speed a… |
@AI21Labs |
Company |
Original |
2025-10-08 |
410,025 |
202 |
28 |
5 |
9 |
78 |
| Congrats @IBM on the release of Granite 4.0! We’re so excited to welcome another Mamba-Transformer model to the mix - a… |
@AI21Labs |
Company |
Original |
2025-10-06 |
1,357 |
19 |
3 |
0 |
1 |
10 |
| 🚀 See you next week at @WorldSummitAI in Amsterdam (Oct 8–9)!📍 Booth G44 – Meet the @AI21Labs team and see how we’re sh… |
@AI21Labs |
Company |
Original |
2025-10-03 |
604 |
4 |
0 |
0 |
0 |
0 |