Rebuilding Among AIs from Six Log Files
Blog post from Hugging Face
A developer independently reconstructed Antim Labs’ Among AIs social-deduction benchmark from six published game logs, deriving a 59×34 map and deterministic game engine that matched all recorded object observations and frames despite lacking the original source code. The rebuilt benchmark ran 90 games among six 27B–35B open models, finding that crewmates won 80% of matches and that models were generally competent at navigation, task execution, spatial-temporal evidence analysis, and identifying contradictions, but struggled to sustain deception as impostors. Many agents displayed limited role-dependent behavior, attempted illegal actions such as impostor task completion or unsupported reports, and made voting decisions that remained wrong in 61% of ejections. A central observation was that some models leaked private deception plans into public chat, including one impostor that described its strategy before being voted out, suggesting that planning and outward communication were not reliably separated under pressure. Results also indicated that verbosity did not predict success, with a concise model tying the highest-scoring, much more verbose model, while the author concludes that agents using models of this size should structurally separate private reasoning from public outputs rather than relying on the model to conceal intent.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 2 | 1,189 | 251 | 109 | -83% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.