Ollama's highest performance on Apple Silicon yet with MLX
Blog post from Ollama
Ollama's updated MLX engine, designed for optimal performance on Apple Silicon, leverages Apple's unified memory and the Metal-backed MLX framework to produce faster, higher-quality responses while using less memory. This update includes support for NVIDIA’s model-optimized NVFP4 format, which enhances output quality compared to other 4-bit quantization formats by closely tracking model weights' local dynamic range and reducing quantization loss. The engine is now 20% faster due to optimizations like fusing operations into single Metal kernels and more efficient GPU-backed sampling. Additionally, Ollama's new snapshot system improves responsiveness in agent workflows by saving model state at key points, allowing for seamless resumption and efficient handling of multi-agent sessions, reasoning models, and conversation branching. This system ensures that only new directions in conversations need processing, conserving memory and enhancing performance. Users can access these improvements by downloading the latest version of Ollama and running models such as Gemma 4 12B with the MLX engine.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| MLX | 12 | 24 | 8 | 5 | +118% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.