POCKET: a 35-billion-parameter model that runs on your iPhone — and on your PC with no GPU
Blog post from Hugging Face
POCKET is a groundbreaking family of on-device builds that enables a 35-billion-parameter language model to run efficiently on devices like iPhones and PCs without needing a GPU. Derived from VIDRAFT's Darwin-36B-Opus, POCKET achieves this by using a sparse Mixture-of-Experts architecture that activates only about 3 billion parameters per token, maintaining high quality and coherence without sacrificing convenience. The model runs faster on both CPU and GPU compared to other leading models, and it's compatible with existing tools like LM Studio and PocketPal. Various builds cater to different devices and RAM capacities, with the performance highly dependent on RAM size rather than GPU presence. POCKET employs domain expert pruning and MoE-aware mixed precision to balance size and performance, especially for language-specific builds like Korean and English. The model's efficiency is showcased by its ability to generate text rapidly, even on devices with limited hardware capabilities, and it is freely available under the Apache-2.0 license for modification and redistribution.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.