BananaMind 2 Pro: We've (almost) matched SmolLM2 at 20x fewer tokens... Trained On a 5070 Ti
Blog post from Hugging Face
Banaxi-Tech reports that its approximately 140-million-parameter BananaMind 2 Pro model was trained on 100 billion tokens using a mixture led by FineWeb-Edu and DCLM, with supplementary synthetic educational, mathematics, and Python data, and that the full run was completed on an RTX 5070 Ti consumer GPU. Compared with SmolLM-135M and SmolLM2-135M, trained on 600 billion and 2 trillion tokens respectively, the model reportedly retains an average 96.08% of SmolLM2’s benchmark performance despite using 20 times fewer tokens, while exceeding the older SmolLM model on several arithmetic, HellaSwag, and Base Bench measures. SmolLM2 remains ahead on most individual benchmarks, notably ARC Easy, INT Index, and code tasks, although BananaMind 2 Pro was ranked ahead of it in a limited SLM Arena user-preference snapshot, with an 81.7 Elo advantage based on 51 versus 43 battles. The author argues that data quality, mixture, architecture, optimization, tokenization, and training strategy can substantially affect small-model performance beyond token count alone, while identifying coding, ARC Easy, and INT Index as priorities for future work and planning new architectures for subsequent BananaMind models.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 4 | 747 | 162 | 79 | -85% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.