Nvidia releases lightweight open-source Nemotron 3.5 Lightning AI model that runs on a single GPU; Nemotron 3.5 Lightning accelerates token generation by 4x but agentic tasks only speed up by 30%

Nvidia released its open-weight Nemotron 3.5 Lightning AI model, which runs on a single GPU and accelerates token generation by 4x. However, agentic tasks see only a 30% speedup as orchestration remains the bottleneck.