18 August 2026
Nvidia releases efficient model with fewer active parameters
- Nvidia released Nemotron 3.5 Lightning, a model designed to run efficiently by activating only 3 billion of its 30 billion total parameters at any given time.
- The model can predict multiple tokens simultaneously, reducing the number of computational steps needed to generate text.
- Nvidia published research on training methods that prevent efficiency loss between training and real-world deployment for large models with this architecture.
How it was covered
Latent Spaceswyx & Alessio
Nemotron 3.5 Lightning exemplifies a shift toward architecture-level efficiency with its 30B MoE design using 3B active parameters, multi-token prediction support, and research on RL for large MoEs showing zero train-infer mismatch.