Home / Companies / Deepinfra / Blog / Post Details
Content Deep Dive

NVIDIA Nemotron 3 Nano 30B API Benchmarks: Latency & Cost

Blog post from Deepinfra

Post Details
Company
Date Published
Author
Deep
Word Count
1,256
Company Posts That Month
34
Language
English
Hacker News Points
-
Post removed?
No
Summary

NVIDIA Nemotron 3 Nano 30B A3B is a large language model developed by NVIDIA, designed for both reasoning and non-reasoning tasks, featuring a hybrid Mamba-Transformer Mixture-of-Experts architecture. With approximately 31.6 billion total parameters, it efficiently uses only 3.2–3.6 billion active parameters per forward pass, offering the reasoning capabilities of a larger model but with the speed and cost efficiency of a lightweight architecture. Trained on 25 trillion tokens across multiple languages and programming languages, the model can toggle between reasoning modes, optimizing for either direct answers or detailed reasoning traces. DeepInfra, the exclusive API provider for this model, offers competitive pricing with a blended cost of $0.09 per million tokens, and supports both JSON Mode and Function Calling, making it suitable for structured output workflows and agentic AI applications. The deployment boasts a 262k token context window and achieves a 93.7 tokens per second output speed, with a sub-half-second TTFT, making it ideal for real-time applications that require immediate responsiveness.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Agents 4 4,430 1,100 236 -3%
LLM 3 5,932 1,046 223 -2%
Real-time 2 6,296 1,346 246 -2%
AI Guardrails 1 362 123 45 +1%
RAG 1 941 216 85 -48%
Vector Search 1 1,739 413 146 -27%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.