Model-Preserving Adaptive Rounding with YAQA
Blog post from Together AI
YAQA (Yet Another Quantization Algorithm) is a new weight-only LLM post-training quantization method that quantizes models to directly preserve the original model's outputs. YAQA achieves state-of-the-art performance on downstream tasks by reducing the KL divergence to the original model by over 30% compared to existing rounding algorithms. It uses a near-optimal Kronecker-factored approximation of each linear layer's Hessian with respect to the KL, which is then used to quantize models with theoretical guarantees. YAQA has been shown to outperform existing methods in experiments, including reducing the cost of training by 20%, increasing network compression by 117x, and achieving faster training times by 4x.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Model Fine-tuning | 5 | 508 | 150 | 76 | -36% |
| LLM | 3 | 4,437 | 679 | 217 | -3% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.