Home / Companies / Ollama / Blog / Post Details
Content Deep Dive

Ollama's highest performance on Apple Silicon yet with MLX

Blog post from Ollama

Post Details
Company
Date Published
Author
-
Word Count
639
Company Posts That Month
4
Language
-
Hacker News Points
-
Post removed?
No
Summary

Ollama's updated MLX engine, designed for optimal performance on Apple Silicon, leverages Apple's unified memory and the Metal-backed MLX framework to produce faster, higher-quality responses while using less memory. This update includes support for NVIDIA’s model-optimized NVFP4 format, which enhances output quality compared to other 4-bit quantization formats by closely tracking model weights' local dynamic range and reducing quantization loss. The engine is now 20% faster due to optimizations like fusing operations into single Metal kernels and more efficient GPU-backed sampling. Additionally, Ollama's new snapshot system improves responsiveness in agent workflows by saving model state at key points, allowing for seamless resumption and efficient handling of multi-agent sessions, reasoning models, and conversation branching. This system ensures that only new directions in conversations need processing, conserving memory and enhancing performance. Users can access these improvements by downloading the latest version of Ollama and running models such as Gemma 4 12B with the MLX engine.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
MLX 12 24 8 5 +118%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.