Home / Companies / Ollama / Blog / October 2025

October 2025 Summaries

6 posts from Ollama

Filter
Month: Year:
Post Summaries Back to Blog
Ollama, in collaboration with OpenAI and ROOST, is introducing gpt-oss-safeguard reasoning models to enhance safety classification tasks. These models, available in two sizes (20B and 120B) and licensed under Apache 2.0, are designed to reason about safety, enabling use cases like filtering and content labeling. The models utilize a "bring your own policy" approach, allowing customization to various products and use cases with minimal engineering. They emphasize reasoned decisions for greater transparency and trust in policy application, and offer configurable reasoning effort to match specific needs. Evaluated against internal and external datasets, including a moderation dataset and the ToxicChat benchmark, these models offer organizations the flexibility to study, modify, and deploy critical safety technologies. ROOST, established in 2025, is a non-profit dedicated to providing open-source safety tools, underscoring the collaborative effort to advance online safety solutions.
Oct 29, 2025 457 words in the original blog post.
MiniMax M2, now accessible on Ollama’s cloud, is a model optimized for coding and agentic workflows, demonstrating superior intelligence in areas like mathematics, science, and coding according to Artificial Analysis benchmarks. It excels in developer tasks such as multi-file edits and test-validated repairs, showing strong performance in coding environments like terminals and IDEs. The model is capable of executing complex toolchains and efficiently recovers from challenges, with its 10 billion activated parameters offering low latency and high throughput. MiniMax M2 can be integrated with tools like VS Code, Zed, and Droid, and accessed via Ollama’s cloud API, making it highly deployable for both interactive agents and batch processing tasks.
Oct 28, 2025 460 words in the original blog post.
Performance tests were conducted on the NVIDIA DGX Spark using the latest firmware and Ollama version 0.12.6 to evaluate its capabilities with various models, including OpenAI's gpt-oss and other models like gemma3, llama3.1, and deepseek-r1. The tests involved generating a summary of "A Tale of Two Cities" with a constraint of 500 tokens, caching disabled, and temperatures set to zero, with each test repeated ten times. The results showed variations in token processing speeds depending on model size and quantization levels, with the gpt-oss models provided by OpenAI being tested through Ollama, which retains intended BF16 attention layers. The DGX Spark firmware can be updated via the DGX Dashboard or CLI, requiring Ubuntu distribution upgrades, and Ollama can be installed for running models. Additionally, OpenAI's Codex can be installed and used alongside Ollama for seamless integration, with the DGX Spark supporting large models like gpt-oss-120b, thanks to its substantial VRAM capacity.
Oct 23, 2025 400 words in the original blog post.
Ollama's cloud service provides access to advanced models like GLM-4.6 and Qwen3-Coder-480B, with Qwen3-Coder-30B now optimized for faster tool calling. Users can run these models in the cloud or locally if they have over 300GB of VRAM. The service offers integrations with popular coding tools such as VS Code, Zed, and Droid, allowing users to select and manage models through easy-to-use interfaces. Additionally, Ollama provides an API for cloud model access, enabling developers to integrate AI functionalities into their applications by setting up an API key and utilizing API endpoints. Detailed instructions for using these models across different platforms and tools are included in Ollama's documentation, catering to a range of coding environments and developer needs.
Oct 16, 2025 439 words in the original blog post.
Qwen3-VL, the latest and most advanced vision language model in the Qwen series, is now accessible via Ollama’s cloud platform, with local availability planned for the future. This model boasts a range of capabilities, including visual agent functions for operating GUIs, visual coding boost for generating code from images or videos, and advanced spatial perception for 2D and 3D grounding. It supports long context and video understanding with a native 256K context expandable to 1M, enhanced multimodal reasoning particularly in STEM fields, and upgraded visual recognition of diverse objects and languages. Additionally, it features expanded OCR capabilities supporting 32 languages and improved text understanding that aligns with pure language models. Users can interact with the model using Ollama’s CLI, API, and JavaScript/Python libraries, and Ollama offers OpenAI-compatible API endpoints for seamless integration.
Oct 14, 2025 547 words in the original blog post.
NVIDIA has introduced the DGX Spark, a high-performance computing system developed in collaboration with Ollama, designed to efficiently run language models out-of-the-box. Equipped with the NVIDIA GB10 Grace Blackwell Superchip, the DGX Spark offers 1 petaFLOP of performance and features 128GB of memory, supporting a range of models from companies like Alibaba, Meta, and OpenAI, among others, in Ollama’s library. Users can also upload custom or fine-tuned models for enhanced flexibility. NVIDIA and Ollama are actively optimizing the system’s performance for various applications, including chat, document processing, coding, and multimodal workflows, inviting users to explore its capabilities and potential uses.
Oct 13, 2025 145 words in the original blog post.