Introducing NVIDIA Nemotron 3.5 Lightning
Blog post from Baseten
NVIDIA Nemotron 3.5 Lightning is a 30-billion-parameter mixture-of-experts model with 3 billion active parameters, distilled from Nemotron 3 Ultra and designed for high-volume, always-on agentic applications such as personal assistance, financial services, cybersecurity, telecom, and retail. Available through Baseten Dedicated Inference on NVIDIA infrastructure, it supports a one-million-token context window and is presented as delivering nearly four times the throughput of similar open models while reducing task-completion time by 30% through faster reasoning and token generation. Baseten and CodeRabbit tested the model for high-volume code-review routing, using supervised fine-tuning and reinforcement learning to improve route agreement and reliability. The resulting rank-16 LoRA adapter reportedly achieved about 4% higher accuracy than the baseline, used roughly half the API cost, generated 63.4% fewer output tokens, and reached 314.82 aggregate output tokens per second across eight concurrent requests on an A100 GPU.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Model Fine-tuning | 3 | 278 | 80 | 43 | -70% |
| Reinforcement learning | 1 | 43 | 19 | 12 | -56% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.