NVIDIA Puts Groq 3 LPX Into Full Production for Agentic AI Inference
NVIDIA says Groq 3 LPX is now in full production as an interactive inference accelerator for Vera Rubin, targeting extremely fast token generation for long-context and multi-step AI agents.
What NVIDIA announced
On August 24, 2026, NVIDIA announced that Groq 3 LPX is in full production. The accelerator extends the NVIDIA Vera Rubin platform and is designed specifically for the token-generation side of latency-sensitive agentic inference. NVIDIA says Nebius is the first AI cloud provider adopting the system.
Why LPX exists
Agentic workloads have two different compute pressures: processing large contexts and generating each next token quickly enough for an agent to remain responsive while it plans, calls tools, reads results and iterates. NVIDIA positions Groq 3 LPX as a specialized complement to Vera Rubin NVL72 for the second problem, with the platform able to split or co-execute inference work across Rubin GPUs and LPX accelerators.
Performance claim and architecture
NVIDIA reports that Artificial Analysis measured 3,431 output tokens per second on a 100K-context benchmark using Gemma 4 31B. NVIDIA's technical description attributes the low latency to deterministic compiler scheduling, fine-grained overlap of communication and compute, SRAM-based execution and tightly planned chip-to-chip transfers. The company also describes configurations for prefill/decode disaggregation, attention/FFN disaggregation and speculative external-drafter decoding.
Why it matters
The launch reflects a broader shift in AI infrastructure from training-only optimization toward inference systems tuned for interactive agents. If specialized decode accelerators can keep long-running agents responsive while GPUs handle large-context and high-throughput work, AI clouds may increasingly deploy heterogeneous inference stacks instead of relying on one accelerator type for every stage.
Availability
NVIDIA describes Groq 3 LPX as in full production, not a future preview. Actual service availability, regional rollout and cloud capacity still depend on NVIDIA systems and adoption by infrastructure providers such as Nebius.
This article is built from the source material below. Open the originals for full context and the latest updates.