NVIDIA Nemotron 3.5 Lightning Is Live on DeepInfra: Day-Zero Access to the Fastest Open Model for Agents
Blog post from Deepinfra
DeepInfra has launched day-zero serverless API access to NVIDIA Nemotron 3.5 Lightning, an open 30-billion-parameter hybrid Mixture-of-Experts model designed for high-volume, always-on AI agents. Activating 3 billion parameters per token, the model supports up to 1 million tokens of context and uses DFlash speculative decoding and multi-token prediction to accelerate structured outputs such as tool calls and multi-step plans. NVIDIA and DeepInfra claim it delivers up to four times the throughput of comparable open models and up to 30% faster agent task completion, with reported benchmark scores of 86.5% on PinchBench agent productivity and 69.9% on AA-Omniscience Non-Hallucination. It is positioned as a customizable, high-throughput model for specialized workflows in areas including personal assistance, finance, cybersecurity, telecom, and retail, potentially alongside model-routing systems such as NVIDIA NeMo Switchyard. The OpenAI-compatible endpoint supports streaming, tool calling, and JSON mode, costs $0.05 per million input tokens and $0.20 per million output tokens, and DeepInfra states that it does not retain requests or use customer data for model training.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Real-time | 3 | 4,432 | 1,050 | 222 | -31% |
| Serverless | 2 | 783 | 217 | 99 | +1% |
| AI Agents | 1 | 5,780 | 1,243 | 245 | -15% |
| LLM | 1 | 5,068 | 1,020 | 229 | -34% |
| Vector Search | 1 | 2,358 | 371 | 127 | +5% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.