NVIDIA Nemotron 3 Super now available on Nebius Token Factory
Blog post from Nebius
NVIDIA Nemotron 3 Super, now accessible on Nebius Token Factory, is a 120 billion parameter hybrid Mixture of Experts (MoE) model designed for multi-agent applications and complex reasoning tasks, featuring 12 billion active parameters per inference step and capable of handling up to 1 million token context length. Optimized for agentic systems, it utilizes a hybrid Transformer–Mamba architecture with MoE routing to enhance compute efficiency while maintaining high reasoning performance. The model is suited for various production use cases, including software development workflows, deep research agents, financial document processing, and cybersecurity analysis. Offering open weights, datasets, and training recipes, it supports multi-token prediction for expedited long-form generation. Nemotron 3 Super can be deployed on Nebius Token Factory through dedicated GPU endpoints with autoscaling throughput and OpenAI-compatible API integration, providing options for EU or US deployment with optional zero-retention inference. This setup allows teams to transition from model access to production deployment without managing GPU clusters, available for deployment via API or testing in the Playground.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.