Same bytes, closer to the original: two lines of AutoRound we had wrong
Blog post from Hugging Face
Archsloth reports that matching AutoRound’s optimization scheme to the exported GGUF format and enabling its sign-gradient rounding search substantially reduced KL-divergence drift versus comparable Unsloth quantizations while preserving identical file sizes, tensor-type layouts, and inference speed. Across a Qwen3-4B Q4_K_M comparison, the reported improvements ranged from about 21% to 54% on ten language and code evaluation axes, with smaller but still positive gains at higher bit widths and on selected 9B and 27B vision-language models. The authors argue that calibration data is especially consequential for AutoRound because it directly influences rounding directions, unlike importance-matrix calibration, and find that multilingual sample-level interleaving and inclusion of code can create measurable quality tradeoffs across domains. They also report unsuccessful approaches, including larger calibration sets, certain sequence lengths, healing, selective bit allocation, mixed precision, and Q5_K exports affected by a packing deviation, while declining to publish configurations that lost on their selected measures. The post emphasizes reproducibility through released model files, corpora, evaluation materials, raw logs, build commands, and comparisons based primarily on KL divergence rather than perplexity, which it argues is less sensitive to output-distribution changes caused by quantization.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.