NVIDIA’s Nemotron 3 Super just cracked the code on real agentic scale: 120B open MoE (12B active), 1M-token context killing context explosion + thinking tax, hybrid Mamba-Transformer + Latent MoE + multi-token prediction for 5x throughput and 2x accuracy, topping DeepResearch Bench while fully open with 10T+ tokens.
We’re already integrating it to push our autonomous research agents into uncharted territory.
Game on - who’s ready to co-build the future?