LFM2.5 Q4_0 Checkpoints from Quantization-Aware Distillation
Blog post from Hugging Face
Liquid AI has released Quantization-Aware Distillation (QAD) Q4_0 GGUF checkpoints for its LFM2.5 230M, 350M, 1.2B-Instruct, and 2.6B models, aiming to preserve model quality while retaining the low memory use and speed of standard 4-bit quantization. QAD distills a high-precision teacher into a quantized student, recovering an average of 97% of the BF16 accuracy normally lost through quantization across benchmarks covering reasoning, instruction following, tool use, agentic tasks, and math. Tests on a MacBook Pro, NucBox EVO-X2, Samsung Galaxy S26 Ultra, and Raspberry Pi 5 found that the smaller QAD models matched Q5_K_M quality with 4–33% greater decoding throughput, while the larger models matched Q4_K_M quality with 3–14% higher throughput. The checkpoints are available on Hugging Face and can be run through llama.cpp or other runtimes supporting GGUF Q4_0 files.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.