NVIDIA Nemotron 3.5 Lightning, an open model designed for high-volume agentic workloads, is now available in Amazon SageMaker JumpStart. According to AWS, the model uses a 30B Mixture-of-Experts architecture with 3B active parameters. The blog post outlines how to deploy the model and states it delivers up to 4x higher throughput and up to 30% faster task completion for always-on agents.

Why it matters

Mixture-of-Experts designs activate only a subset of parameters per request, which can support the throughput and latency requirements of continuously running agent systems. Availability in a managed catalog like SageMaker JumpStart lowers the setup effort for deploying the model.

Who should care

Developers and teams building agentic applications on AWS who are evaluating open models for high-volume, always-on workloads.