Nucleus-Image: Scaling Text-to-Image with Sparse Mixture of Experts
Blog post from Hugging Face
Nucleus-Image is a groundbreaking 17-billion-parameter text-to-image diffusion model developed by Nucleus AI, which utilizes a sparse mixture-of-experts (MoE) approach that activates only about 2 billion parameters per forward pass. This innovative design allows the model to match or outperform competitors like Qwen-Image and Imagen 4 across various benchmarks such as GenEval, DPG-Bench, and OneIG-Bench, all achieved through pre-training without reinforcement learning or human preference tuning. By decoupling the capacity from compute, Nucleus-Image offers the vast knowledge of a larger network with the efficiency of a smaller one, making it the first fully open-source MoE diffusion model of its quality. The model's architecture introduces several innovations, including decoupled routing for stable MoE diffusion, text tokens serving only as keys and values to optimize performance, and progressive sparsification tied to resolution. Additionally, the model leverages custom Triton kernels and expert parallelism to enhance computational efficiency, training on a meticulously curated dataset of 700 million images and 1.5 billion captions. Nucleus-Image is designed to be a foundational model for the community, providing open access to its weights, training code, and dataset recipe, with future developments focusing on higher-resolution variants and further optimizations.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 2 | 6,889 | 1,263 | 265 | -9% |
| Reinforcement learning | 1 | 109 | 54 | 27 | -40% |
| Vector Search | 1 | 1,977 | 499 | 171 | -39% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.