Up to 3.2x Faster Inference with LFM2.5-DSpark
Blog post from Hugging Face
Liquid AI has released DSpark draft-model checkpoints for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B, using speculative decoding to accelerate token generation while preserving identical greedy-decoding outputs and benchmark accuracy. DSpark combines a parallel draft backbone, a sequential Markov-chain head to improve later-token acceptance, and a confidence-based verifier that discards candidate suffixes when verification is inefficient; the draft models contain roughly 296–328 million parameters and were trained on mixed instruction, chat, code, and function-calling data. Tests using SGLang on an Nvidia H100 GPU and llama.cpp with Metal on an M4 Max MacBook Pro showed average GPU speedups of 2.10x to 2.67x across the models, with peak performance reaching 3.18x for the 8B-A1B model, while on-device results ranged from an average 1.18x for the mixture-of-experts model to 2.54x for the 1.2B model. The company also reports that DSpark reduced LFM2.5-2.6B function-calling latency by an average of 57% in multi-tool scenarios, although performance varies with token acceptance rates and current hardware backend limitations. The checkpoints are available in Safetensors and GGUF formats, with upstream integration support for SGLang and llama.cpp.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.