developer.nvidia.com
2026-08-11
NVIDIA has launched Nemotron 3.5 Lightning, a specialized 30B mixture-of-experts (MoE) model designed for high-volume, low-latency execution in long-running AI agents. Built for efficiency, it features only 3B active parameters, enabling performance comparable to larger models at a fraction of the c
developer.nvidia.com
2026-06-04
NVIDIA introduces Nemotron 3 Ultra, a new open model designed to accelerate reasoning and efficiency for long-running agents. As conversational AI evolves into complex multi-turn systems, token accumulation increases costs and risks goal drift. Nemotron 3 Ultra addresses this with architectural inno