Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

Trained 210M text-to-image model from scratch on one GPU: what actually mattered

Blog post from Hugging Face

Post Details
Company
Date Published
Author
Ivan Mikhnenkov
Word Count
1,731
Company Posts That Month
82
Language
-
Hacker News Points
-
Post removed?
No
Summary

Ivan Mikhnenkov describes training tinydit, a 210 million-parameter text-to-image diffusion transformer from scratch in 3.5 days on one RTX PRO 6000 using 4.2 million images, while relying on frozen FLUX.2 autoencoder and Flan-T5 text-encoder components. The project found that carefully curated, mixed-source data, paired long and short captions, limited synthetic imagery, and aspect-ratio bucketing were more consequential than many architectural changes, avoiding the losses caused by square cropping. The model uses a contemporary DiT architecture with cross-attention, 2D rotary embeddings, QK normalization, SwiGLU, adaptive layer normalization, and learned register and null-attention tokens to manage global state. Its training recipe combines rectified flow with a noise-timestep shift tailored to 32-channel latents, auxiliary directional and representation-dispersion losses, EMA weights, and compiled training, which reportedly doubled speed while reducing memory use. Training loss was considered useful for detecting instability or overfitting but insufficient for assessing image quality, so progress was monitored with FID, DINOv2-based metrics, object accuracy, and preference models. Tinydit performs reliably on individual objects, animals, scenes, colors, and simple spatial prompts but struggles with text, faces, crowds, precise counts, clocks, and complex geometry; future work will use reinforcement learning to improve these limitations.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Data Pipeline 1 34 23 18 -90%
Reinforcement learning 1 17 7 5 -82%
Vector Search 1 265 57 33 -89%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.