NVIDIA launches Nemotron 3.5 Lightning, a sparse open model built for fast agents

NVIDIA has released Nemotron 3.5 Lightning, an open 30-billion-parameter mixture-of-experts model that activates 3 billion parameters and is aimed at always-on agents handling large volumes of specialized work. NVIDIA says it can produce output up to four times faster than similar-sized models, though the announcement does not identify the comparison models, hardware or serving setup. The attached Artificial Analysis chart places Lightning at 24 on its composite Intelligence Index, level with gpt-oss-120b (high) in that snapshot and below several larger or competing models. The release matters because it targets a practical agent trade-off: enough capability for repeated specialized tasks with a much smaller active compute footprint, while leaving the speed claim dependent on deployment conditions.