Accelerating on-device AI: A look at Arm and Google AI Edge optimization
Blog post from Google Cloud
AI technology is advancing towards multimodal capabilities that include on-device image and audio generation, allowing developers to create personalized consumer experiences. Traditionally, executing large AI models at the edge has involved a tradeoff between high latency on CPUs and using specialized, fragmented accelerators. The Arm Scalable Matrix Extension 2 (SME2) resolves this by integrating matrix-compute units into the CPU, enhancing its performance as an AI accelerator and improving inference speeds for generative AI tasks by up to 5x. Google's AI Edge platform further simplifies AI deployment on Arm hardware, supporting automatic runtime optimizations through tools like LiteRT, XNNPACK, and Arm KleidiAI, which enhance efficiency by targeting math-intensive kernels. By leveraging this integration, developers can transform models like Stability AI's stable-audio-open-small into optimized, mixed-precision implementations suitable for high-performance edge deployment, while Google's AI Edge Quantizer and Model Explorer facilitate model compression and performance optimization. This synergy enables significant performance improvements, reducing audio generation time and memory usage while maintaining audio quality, opening opportunities for scaling applications across a wide range of CPU-powered devices globally.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Local AI | 2 | 47 | 28 | 21 | -27% |
| LLM | 1 | 9,074 | 1,640 | 224 | +53% |
| Vector Search | 1 | 2,268 | 422 | 128 | +30% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.