Trained 210M text-to-image model from scratch on one GPU: what actually mattered
Blog post from Hugging Face
Ivan Mikhnenkov describes training tinydit, a 210 million-parameter text-to-image diffusion transformer from scratch in 3.5 days on one RTX PRO 6000 using 4.2 million images, while relying on frozen FLUX.2 autoencoder and Flan-T5 text-encoder components. The project found that carefully curated, mixed-source data, paired long and short captions, limited synthetic imagery, and aspect-ratio bucketing were more consequential than many architectural changes, avoiding the losses caused by square cropping. The model uses a contemporary DiT architecture with cross-attention, 2D rotary embeddings, QK normalization, SwiGLU, adaptive layer normalization, and learned register and null-attention tokens to manage global state. Its training recipe combines rectified flow with a noise-timestep shift tailored to 32-channel latents, auxiliary directional and representation-dispersion losses, EMA weights, and compiled training, which reportedly doubled speed while reducing memory use. Training loss was considered useful for detecting instability or overfitting but insufficient for assessing image quality, so progress was monitored with FID, DINOv2-based metrics, object accuracy, and preference models. Tinydit performs reliably on individual objects, animals, scenes, colors, and simple spatial prompts but struggles with text, faces, crowds, precise counts, clocks, and complex geometry; future work will use reinforcement learning to improve these limitations.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Data Pipeline | 1 | 34 | 23 | 18 | -90% |
| Reinforcement learning | 1 | 17 | 7 | 5 | -82% |
| Vector Search | 1 | 265 | 57 | 33 | -89% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.