Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

BananaMind 2 Pro: We've (almost) matched SmolLM2 at 20x fewer tokens... Trained On a 5070 Ti

Blog post from Hugging Face

Post Details
Company
Date Published
Author
Banaxi
Word Count
1,568
Company Posts That Month
39
Language
-
Hacker News Points
-
Post removed?
No
Summary

Banaxi-Tech reports that its approximately 140-million-parameter BananaMind 2 Pro model was trained on 100 billion tokens using a mixture led by FineWeb-Edu and DCLM, with supplementary synthetic educational, mathematics, and Python data, and that the full run was completed on an RTX 5070 Ti consumer GPU. Compared with SmolLM-135M and SmolLM2-135M, trained on 600 billion and 2 trillion tokens respectively, the model reportedly retains an average 96.08% of SmolLM2’s benchmark performance despite using 20 times fewer tokens, while exceeding the older SmolLM model on several arithmetic, HellaSwag, and Base Bench measures. SmolLM2 remains ahead on most individual benchmarks, notably ARC Easy, INT Index, and code tasks, although BananaMind 2 Pro was ranked ahead of it in a limited SLM Arena user-preference snapshot, with an 81.7 Elo advantage based on 51 versus 43 battles. The author argues that data quality, mixture, architecture, optimization, tokenization, and training strategy can substantially affect small-model performance beyond token count alone, while identifying coding, ARC Easy, and INT Index as priorities for future work and planning new architectures for subsequent BananaMind models.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 4 747 162 79 -85%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.