NVIDIA Nemotron 3.5 Lightning
Blog post from Ollama
NVIDIA Nemotron 3.5 Lightning is a 30-billion-parameter open model, with 3 billion active parameters per token, now available through Ollama for local deployment on compatible NVIDIA hardware and Apple silicon. Built on a hybrid Mixture-of-Experts architecture, it is designed for long-running agentic workflows involving tool calls, coding, context gathering, retries, and multi-step task completion while keeping local data on the user’s device. Its features include up to a 1 million-token context window, speculative decoding and multi-token prediction for improved inference throughput, and open weights and datasets that allow developers to customize it for specialized tasks. Suggested applications include personal assistants, coding agents, security operations, and hybrid workflows in which routine high-volume tasks run locally while more demanding steps are sent to cloud models through the same API or CLI. NVIDIA reports that the model achieves up to four times higher throughput, 30% faster task completion, and competitive accuracy on agentic, coding, and reasoning benchmarks compared with similarly sized open models.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.