Home / Companies / Google Cloud / Blog / Post Details
Content Deep Dive

On-device GenAI in Chrome, Chromebook Plus, and Pixel Watch with LiteRT-LM

Blog post from Google Cloud

Post Details
Company
Date Published
Author
Yu-hui Chen, and Ram Iyengar
Word Count
1,726
Company Posts That Month
20
Language
English
Hacker News Points
-
Post removed?
No
Summary

Running large language models (LLMs) directly on devices like Chrome, Chromebook Plus, and the Pixel Watch is made possible through LiteRT-LM, a framework designed for efficient and high-performance on-device inference. This approach offers the advantages of offline availability and cost-effectiveness, eliminating per-API-call costs and making LLMs practical for frequent tasks such as text summarization and proofreading. LiteRT-LM addresses the challenges of deploying gigabyte-scale models across various hardware by utilizing a modular, open-source design that supports multiple platforms and accelerators, including CPU, GPU, and NPU. The framework's architecture, consisting of an Engine and Session system, allows shared resources to be managed efficiently while enabling customization through lightweight adapters. This system enhances flexibility and scalability, adapting to different device constraints, from powerful smartphones to resource-limited wearables like the Pixel Watch, where a minimal pipeline can be constructed to optimize binary size and memory usage. The framework also integrates with Google's broader AI Edge stack, supporting developers in building custom LLM-powered applications and scaling them across diverse platforms, highlighting its utility in products like Chrome and the Pixel Watch's Smart Replies feature.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 20 3,636 538 190 -7%
AI Model Fine-tuning 2 276 96 58 -51%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.