Home / Companies / Google Cloud / Blog / Post Details
Content Deep Dive

Large Language Models On-Device with MediaPipe and TensorFlow Lite

Blog post from Google Cloud

Post Details
Company
Date Published
Author
Mark Sherwood, and Juhyun Lee
Word Count
1,525
Company Posts That Month
7
Language
English
Hacker News Points
-
Post removed?
No
Summary

Since its introduction in 2017, TensorFlow Lite has enabled on-device machine learning, with MediaPipe enhancing its capabilities in 2019 by supporting full ML pipelines. The experimental release of the MediaPipe LLM Inference API marks a significant advancement, allowing Large Language Models (LLMs) to function fully on-device, despite their demanding memory and compute requirements. This new API is optimized for web, Android, and iOS platforms, initially supporting models like Gemma, Phi 2, Falcon, and Stable LM. It benefits from various optimizations, including new ops, quantization, and weight sharing, to manage the computational intensity of LLMs. On Android, the API is currently intended for experimental use, with production applications advised to use the Gemini API or Gemini Nano through Android AICore, a feature in Android 14 that supports high-end devices with ML accelerators and safety filters. The API facilitates on-device LLM integration for developers, allowing them to prototype and test LLMs using platform-specific SDKs. Significant optimizations across MediaPipe, TensorFlow Lite, and XNNPack have been implemented to support efficient LLM inference, with strategies like sharing weights, optimizing fully connected operations, and employing custom GPU operations to balance compute and memory demands. The release aims to expand further throughout 2024, with plans for more platforms, models, and broader conversion tools, signaling a transformative step for on-device machine learning.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Model Fine-tuning 1 434 113 72 -8%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.