All stories
AI

NVIDIA Unveils Nemotron-Labs-3-Puzzle-75B-A9B: 2.03x Server Throughput Boost for LLMs

NVIDIA's new compressed hybrid Mixture-of-Experts LLM, Puzzle-75B-A9B, significantly boosts server throughput by 2.03x while maintaining strong accuracy, making advanced AI deployment more efficient.

Source:MarkTechPost·2 min read·Jul 9

This content was summarized and interpreted by AI; it may contain errors — please verify accuracy with the original sources. Learn more

Share

Listen to this story

0:00 / 0:00
NVIDIA Unveils Nemotron-Labs-3-Puzzle-75B-A9B: 2.03x Server Throughput Boost for LLMs

NVIDIA has unveiled Nemotron-Labs-3-Puzzle-75B-A9B, a new compressed hybrid Mixture-of-Experts (MoE) large language model (LLM) that delivers a remarkable 2.03x server throughput increase compared to its predecessor, Nemotron-3-Super, at matched user throughput. This significant efficiency gain is achieved while preserving strong accuracy across a diverse set of benchmarks, including reasoning, coding, multilingual, long-context, and agentic tasks.

The new model, often referred to as Puzzle-75B-A9B, is a compressed variant of the 120.7 billion total parameter Nemotron-3-Super. Through a sophisticated post-training compression framework dubbed "Iterative Puzzle," NVIDIA has reduced the total parameters to 75.3 billion and active parameters from 12.8 billion to 9.3 billion, representing a substantial reduction in model size. The Iterative Puzzle method strategically alternates hardware-aware structural compression with short knowledge distillation recovery phases, ensuring minimal performance degradation. This process includes heterogeneous MoE channel pruning, active expert reduction, and Mamba SSM state size pruning, while maintaining the parent's efficient attention layers.

This breakthrough in model compression holds profound implications for the deployment of advanced AI. The ability to achieve 2.03x higher server throughput, with a potential rise to 4.63x with Multi-Token Prediction (MTP), translates directly into lower operational costs and enhanced scalability for enterprises and cloud service providers. For instance, Puzzle-75B-A9B boosts sustainable 1M-token concurrency on a single H100 GPU from one request to eight, driven by a significant reduction in NVFP4 weights from 70 GB to 44.5 GB. This makes powerful LLMs more accessible and cost-effective for a wider range of interactive, reasoning-heavy, and long-context workloads. NVIDIA’s strategic focus on optimizing inference efficiency without compromising quality signals a pivotal shift towards democratizing high-performance AI, accelerating its integration into real-world applications.

Watch (Shorts)