Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

Extreme Overtraining in Tiny Language Models

Blog post from Hugging Face

Post Details
Company
Date Published
Author
Banaxi
Word Count
832
Company Posts That Month
74
Language
-
Hacker News Points
-
Post removed?
No
Summary

An experiment training a 0.9-million-parameter language model on up to 200 billion tokens found that extreme token-to-parameter ratios can substantially degrade small-model benchmark performance. Using a six-layer architecture, a 384-token vocabulary, Muon and AdamW optimization, and FineWeb-HQ plus Cosmopedia v2 data, the model’s aggregate INT Index peaked at 4.55 after 20 billion tokens, or roughly 22,000 tokens per parameter, before declining to 3.31 by 180 billion tokens, a 27.3% reduction. Most evaluated benchmarks, including PIQA, ARC-Challenge, and HellaSwag, performed worse late in training, while a control run at the conventional Chinchilla ratio of about 20 tokens per parameter was essentially chance-level, indicating severe undertraining. The results suggest that very small models benefit from substantially more data than Chinchilla-optimal scaling predicts, but that performance may peak around 22,000–30,000 tokens per parameter and deteriorate when training continues far beyond that range.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 1 2,312 357 123 +3%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.