On-device GenAI in Chrome, Chromebook Plus, and Pixel Watch with LiteRT-LM
Blog post from Google Cloud
Running large language models (LLMs) directly on devices like Chrome, Chromebook Plus, and the Pixel Watch is made possible through LiteRT-LM, a framework designed for efficient and high-performance on-device inference. This approach offers the advantages of offline availability and cost-effectiveness, eliminating per-API-call costs and making LLMs practical for frequent tasks such as text summarization and proofreading. LiteRT-LM addresses the challenges of deploying gigabyte-scale models across various hardware by utilizing a modular, open-source design that supports multiple platforms and accelerators, including CPU, GPU, and NPU. The framework's architecture, consisting of an Engine and Session system, allows shared resources to be managed efficiently while enabling customization through lightweight adapters. This system enhances flexibility and scalability, adapting to different device constraints, from powerful smartphones to resource-limited wearables like the Pixel Watch, where a minimal pipeline can be constructed to optimize binary size and memory usage. The framework also integrates with Google's broader AI Edge stack, supporting developers in building custom LLM-powered applications and scaling them across diverse platforms, highlighting its utility in products like Chrome and the Pixel Watch's Smart Replies feature.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 20 | 3,636 | 538 | 190 | -7% |
| AI Model Fine-tuning | 2 | 276 | 96 | 58 | -51% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.