Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

Accelerating vision-language models with LFM2.5-VL-DSpark

Blog post from Hugging Face

Post Details
Company
Date Published
Author
xx, Yuri Khrustalev, Leonie Monigatti, and Viviana Márquez
Word Count
1,011
Company Posts That Month
82
Language
-
Hacker News Points
-
Post removed?
No
Summary

Liquid AI has released LFM2.5-VL-DSpark, an experimental speculative-decoding draft model for its 3B-parameter LFM2.5-VL vision-language model, intended to accelerate inference while preserving the target model’s exact greedy-output quality through token verification. The approximately 280M-parameter, four-layer drafter adds 8.9% to the deployed model size, uses hidden states from selected target-model layers to propose blocks of candidate tokens, and was trained on vision-language supervised fine-tuning data with recommended block sizes of eight or nine tokens. Across vision tasks such as visual question answering, captioning, chart analysis, reasoning, and multi-turn conversation, reported decoding speedups reached 3.13x on an Apple M5 Max, 2.14x on an M3 Ultra, and 2.66x on an H100 GPU, while end-to-end improvements were lower because speculative decoding does not accelerate image encoding or prompt prefill. The model is available in Safetensors and GGUF formats and has initial integration support in llama.cpp, MLX-VLM, and SGLang, with instructions provided for launching each implementation.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
MLX 5 1 1 1 -96%
LLM 2 747 162 79 -85%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.