NVIDIA shared new benchmark results using SemiAnalysis's AgentX test, which measures agentic-coding inference performance across models including Kimi K3, MiniMax M3, GLM5.3, Qwen3.5, and DeepSeek V4 Pro. According to the results, NVIDIA's Blackwell-based GB300 NVL72 server delivers 15x higher throughput per megawatt than the previous Hopper-based H200 NVL8 in DeepSeek-v4-PRO 1.6T, along with 10x lower cost per million tokens. With the larger Kimi K3 2.8T model, Blackwell's advantage grew to an 80x increase in throughput per megawatt versus Hopper, sustaining 215 tokens per second per user.

NVIDIA's next-generation Vera Rubin NVL72 platform posted even larger gains over Blackwell in the same DeepSeek-v4-PRO 1.6T test, with a 30x increase in throughput and 35x lower cost per million tokens. Vera Rubin reached roughly 160 tokens per second per user in interactivity and peaked near 280 tokens per second per user, compared to Blackwell's ceiling of below 180 TPS per user. NVIDIA said its DSX MaxLPS power-management technology can allow AI factories to provision up to 40% more GPUs within the same power budget. NVIDIA noted these figures do not yet include Vera CPU performance for tool calling and reflect only part of the broader seven-chip Vera Rubin platform.

At Hot Chips 2026, NVIDIA announced that its Groq 3 LPX AI inference accelerator has entered full production, following earlier production announcements for Vera CPUs and Vera Rubin servers. Groq 3 LPX is designed to extend Vera Rubin NVL72 by accelerating token generation for latency-sensitive agentic workloads. In a demonstration using the Gemma 4 31B model with a 100,000-token context window, a Vera Rubin NVL72 system paired with Groq 3 LPX achieved 3,400 tokens per second in Artificial Analysis testing, which NVIDIA said was the fastest performance ever recorded on that model. NVIDIA said the combination can cut agentic coding tasks from hours to minutes, offering a 4x improvement in response time over the next-best alternative platform. NVIDIA CEO Jensen Huang described Vera Rubin and LPX as extending the performance gains first delivered by Blackwell into the agentic AI era. Cloud provider Nebius said it plans to deploy Groq 3 LPX through its Nebius Token Factory, with its CTO stating that the technology accelerates the token-generation phase of inference without requiring developers to migrate to a new stack.

Separately, NVIDIA and SpaceXAI announced a partnership under which SpaceXAI will adopt NVIDIA's Vera CPUs and Vera Rubin servers for its Agentic AI operations. SpaceXAI's president said Vera provides the CPU performance and memory bandwidth needed for orchestration, code execution, and data processing tasks that support AI agents. SpaceXAI will use Vera Rubin servers to expand infrastructure for its Grok AI system as it scales toward gigawatt-level computing capacity.

The partnership also extends into orbital computing through a Space-1 Vera Rubin module built for space-based AI workloads. NVIDIA said the module offers up to 25 times the AI compute capability of the previous H100 GPU for orbital use, is powered by solar energy, and shares the same architecture as NVIDIA's terrestrial Vera Rubin chips. It is intended to support geospatial intelligence, autonomous operations, and on-orbit analytics, with planned use by companies including Aetherflux, Axiom Space, and Planet Labs. SpaceXAI's first-generation Starmind AI satellite is set to use this system. Elon Musk stated that SpaceX, working with NVIDIA, has designed a space-optimized Vera Rubin NVL72 system planned for launch in the fourth quarter of next year, with significant scale expected by 2028.

NVIDIA executives framed the announcements as part of a broader push to build computing systems for agentic AI that can take real-time action, not just generate responses. An NVIDIA vice president said Vera gives AI agents the CPU performance to execute code, process data, and coordinate tasks in real time, while extending that architecture from large-scale AI factories to orbital computing.