LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge
Blog post from Hugging Face
Liquid AI introduces LFM2.5-VL-3B, a 3.1-billion-parameter vision-language model designed for fast, private deployment on edge devices and local hardware. Built with a SigLIP2 vision encoder and the LFM2.5-2.6B text backbone, it was trained on roughly 34 trillion tokens with expanded visual data, multilingual vocabulary support, supervised distillation, and reinforcement learning. The release emphasizes improved screen and UI understanding, object grounding, multi-image analysis, document and OCR capabilities, and text- and vision-based function calling. Liquid AI reports that the model performs competitively or leads its size class on a range of multimodal benchmarks, especially for grounding, screen comprehension, documents, and real-world image tasks, while also improving instruction following and tool use. It supports common inference frameworks including Transformers, llama.cpp, MLX, vLLM, SGLang, and ONNX, and is reported to run in about 3 GB of memory with performance ranging from mobile-device inference to high-throughput GPU deployment. The model is available through Hugging Face, with browser demos, documentation, and fine-tuning resources.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 5 | 2,482 | 499 | 155 | -67% |
| AI Model Fine-tuning | 2 | 278 | 80 | 43 | -70% |
| MLX | 1 | 13 | 5 | 2 | -59% |
| Real-time | 1 | 2,081 | 529 | 162 | -65% |
| Reinforcement learning | 1 | 43 | 19 | 12 | -56% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.