Extreme Overtraining in Tiny Language Models
Blog post from Hugging Face
An experiment training a 0.9-million-parameter language model on up to 200 billion tokens found that extreme token-to-parameter ratios can substantially degrade small-model benchmark performance. Using a six-layer architecture, a 384-token vocabulary, Muon and AdamW optimization, and FineWeb-HQ plus Cosmopedia v2 data, the model’s aggregate INT Index peaked at 4.55 after 20 billion tokens, or roughly 22,000 tokens per parameter, before declining to 3.31 by 180 billion tokens, a 27.3% reduction. Most evaluated benchmarks, including PIQA, ARC-Challenge, and HellaSwag, performed worse late in training, while a control run at the conventional Chinchilla ratio of about 20 tokens per parameter was essentially chance-level, indicating severe undertraining. The results suggest that very small models benefit from substantially more data than Chinchilla-optimal scaling predicts, but that performance may peak around 22,000–30,000 tokens per parameter and deteriorate when training continues far beyond that range.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Vector Search | 1 | 2,312 | 357 | 123 | +3% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.