August 2026 Summaries
4 posts from Ollama
Filter
Month:
Year:
Post Summaries
Back to Blog
Ollama has introduced transparent per-token pricing for its Pro, Max, and Team plans, replacing GPU-time-based billing with monthly usage-credit pools that renew each month and can be exceeded at the same published token rates without service fees or hidden limits. Pro costs $20 per month with $60 in included usage, Max costs $100 with $300, and the new Team plan costs $500 with $1,000 in shared usage for unlimited users; free users receive limited monthly access to starter models and can add pay-as-you-go credits for broader access. All plans provide access to current open models, API and coding-agent compatibility, dedicated compute hosted primarily in the US and Europe, zero data retention, and request-level cost visibility, while existing subscribers may retain their current plans or switch to the new model.
Aug 31, 2026
664 words in the original blog post.
Ollama now supports integration with Claude Desktop as a third-party gateway, allowing developers to access Ollama’s local open models and cloud-hosted models while retaining the ability to switch back to Anthropic’s Claude models. Users can enable the connection by opening Ollama, selecting and turning on Claude, after which Ollama automatically configures the gateway; disabling it restores the prior setup. Ollama states that telemetry is disabled by default and that its services follow a Zero Data Retention policy, so prompts and data are not sent to Anthropic or retained by Ollama. Local models can also be configured through Ollama’s Settings page.
Aug 25, 2026
193 words in the original blog post.
NVIDIA Nemotron 3.5 Lightning is a 30-billion-parameter open model, with 3 billion active parameters per token, now available through Ollama for local deployment on compatible NVIDIA hardware and Apple silicon. Built on a hybrid Mixture-of-Experts architecture, it is designed for long-running agentic workflows involving tool calls, coding, context gathering, retries, and multi-step task completion while keeping local data on the user’s device. Its features include up to a 1 million-token context window, speculative decoding and multi-token prediction for improved inference throughput, and open weights and datasets that allow developers to customize it for specialized tasks. Suggested applications include personal assistants, coding agents, security operations, and hybrid workflows in which routine high-volume tasks run locally while more demanding steps are sent to cloud models through the same API or CLI. NVIDIA reports that the model achieves up to four times higher throughput, 30% faster task completion, and competitive accuracy on agentic, coding, and reasoning benchmarks compared with similarly sized open models.
Aug 11, 2026
517 words in the original blog post.
Meta’s Muse Glimmer, the first open model released by Meta Superintelligence Labs, is now available through Ollama with initial support on Apple Silicon via its MLX engine. The 30-billion-parameter multimodal model is licensed under Apache 2.0, supports context lengths exceeding 128K tokens, and is designed for local agent workloads including coding assistants and long-running personal assistants. Ollama integration enables use with tools such as Claude Code, Codex, Pi, OpenClaw, Hermes, OpenCode, and GitHub Copilot, while offering adjustable reasoning levels from low to xhigh to balance speed and task complexity. MLX support adds DFlash acceleration, which reportedly makes Muse Glimmer 1.5 to 1.8 times faster on Apple Silicon, as well as image input capabilities enabled by its 1.8-billion-parameter perception encoder for tasks involving mockups, screenshots, documents, receipts, and charts. Broader platform support and further optimizations for NVIDIA, AMD, and other hardware are planned.
Aug 10, 2026
349 words in the original blog post.