Home / Companies / Together AI / Blog / Post Details
Content Deep Dive

CocktailSGD: Fine-tuning foundation models over 500Mbps networks

Blog post from Together AI

Post Details
Company
Date Published
Author
Jue Wang, Binhang Yuan, Luka Rimanic, Yongjun He, Tri Dao, Beidi Chen, Christopher Re, Ce Zhang
Word Count
234
Company Posts That Month
4
Language
English
Hacker News Points
-
Post removed?
No
Summary

CocktailSGD is a novel communication-efficient training framework designed to train large language models (LLMs) over slow networks, such as 500Mbps connections. This approach combines three distinct compression techniques - random sparsification, top-K sparsification, and quantization - to achieve much greater compression than individual techniques alone. Theoretical analysis justifies the benefit of this hybrid approach, while empirical results show that CocktailSGD achieves up to 117x compression in fine-tuning LLMs without compromising convergence. On a slow network, CocktailSGD only incurs a small slowdown compared to data center networks. The RedPajama-V2 Dataset is conceptualized as a foundation for creating high-quality datasets, and its use requires filtering out data using quality signals that accompany it, depending on the intended application.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Model Fine-tuning 3 138 57 30 -23%
LLM 2 805 142 68 -5%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.