Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

Training a 2.7B MoE from scratch for $200, one GPU at a time

Blog post from Hugging Face

Post Details
Company
Date Published
Author
V S
Word Count
1,299
Company Posts That Month
73
Language
-
Hacker News Points
-
Post removed?
No
Summary

A group of contributors successfully trained a 2.7-billion-parameter Mixture-of-Experts model, NanoColibri-Instruct, from scratch using rented GPUs for a cost of approximately $180-$260, employing a relay training method where contributors sequentially used one GPU at a time. This innovative approach allowed the model to outperform similar-sized dense models in multiple zero-shot tasks with fewer active parameters. Notably, the training was conducted without the use of clusters, and the process, fully documented and reproducible, involved a unique leasing system to ensure that only one contributor trained at a time. The project aimed to demonstrate the feasibility of using small MoE models efficiently on consumer-grade hardware by streaming experts from storage rather than relying on RAM, paving the way for developing larger models like Colibri-Micro and Colibri-Grande. The team learned valuable lessons about cache management and training schedules, which will inform future projects. The open-source nature of this work invites further collaboration and sponsorship opportunities for future developments.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Real-time 4 5,674 1,350 233 -6%
LLM 1 7,115 1,261 236 +13%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.