Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

VisionPsy-Nano: State-of-the-Art On-Device Vision-Language Models

Blog post from Hugging Face

Post Details
Company
Date Published
Author
Khurram Azeem Hashmi, Mohammadreza Zolfaghari, Changdae Park, Rishabh Jain, Nicholas Moratelli, pengfei, Louis, Tianchi Liu, Mathias Buus, and Amril Nurman
Word Count
4,467
Company Posts That Month
73
Language
-
Hacker News Points
-
Post removed?
No
Summary

Tether AI Research has introduced VisionPsy-Nano, a state-of-the-art family of compact vision-language models (VLMs) designed for on-device and edge deployment, offering significant advancements in multimodal understanding. The VisionPsy-Nano family consists of two variants: VisionPsy-Nano-460M, which prioritizes quality, and VisionPsy-Nano-460M-Flash, which is optimized for latency, both achieving high performance across four key capability areas—document understanding & OCR, visual perception, reasoning & knowledge, and instruction following & reliability. VisionPsy-Nano-460M outperforms all models in its ~0.5B parameter class on 16 out of 17 benchmarks, demonstrating superior capabilities in reasoning and knowledge with a notable margin. The Flash variant offers rapid processing on mobile devices, achieving significantly lower time-to-first-token while maintaining close to full model quality. These models are publicly released under Apache 2.0 with open weights, allowing researchers to reproduce benchmark results and enabling application in latency- and memory-constrained environments like smartphones.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 11 7,115 1,261 236 +13%
AI Model Fine-tuning 3 896 206 76 +18%
Reinforcement learning 2 98 52 31 +23%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.