Training a 2.7B MoE from scratch for $200, one GPU at a time
Blog post from Hugging Face
A group of contributors successfully trained a 2.7-billion-parameter Mixture-of-Experts model, NanoColibri-Instruct, from scratch using rented GPUs for a cost of approximately $180-$260, employing a relay training method where contributors sequentially used one GPU at a time. This innovative approach allowed the model to outperform similar-sized dense models in multiple zero-shot tasks with fewer active parameters. Notably, the training was conducted without the use of clusters, and the process, fully documented and reproducible, involved a unique leasing system to ensure that only one contributor trained at a time. The project aimed to demonstrate the feasibility of using small MoE models efficiently on consumer-grade hardware by streaming experts from storage rather than relying on RAM, paving the way for developing larger models like Colibri-Micro and Colibri-Grande. The team learned valuable lessons about cache management and training schedules, which will inform future projects. The open-source nature of this work invites further collaboration and sponsorship opportunities for future developments.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.