NVIDIA Enters Full Production of Groq 3 LPX AI Inference Accelerator Chips, Supercharging Vera Rubin With The Fastest Token Generation Speeds Ever Recorded
NVIDIA announced full production of its Groq 3 LPX AI inference accelerator chips, which boost token generation speeds for its Vera Rubin platforms. The chips enable 3,400 tokens per second, a record for the Gemma 4 31B model, and are being adopted by AI cloud providers like Nebius for advanced infrastructure.
How this was made

The 30-second read
Why it matters
The product launch provides a tangible performance improvement, reinforcing NVIDIA's premium pricing power in AI hardware.
Market read
NVIDIA's new accelerator could drive short-term stock momentum and influence the broader AI hardware market.
What to watch
Potential supply chain constraints or higher-than-expected production costs could limit the accelerator's impact.
Background
NVIDIA's AI roadmap includes the Vera Rubin platform and new inference accelerators aimed at agentic AI workloads.
Ticker impact
NVIDIA announced full production of its Groq 3 LPX AI inference accelerator, delivering 3,400 tokens per second on the Vera Rubin platform.
Potential short-term upside as investors price in the product launch and related revenue opportunities.
First‑time disclosure of a high‑performance AI chip with concrete performance metrics; aligns with NVIDIA's AI growth narrative.
Market effects
Strengthens the AI hardware sector and may pressure competitors to accelerate their own accelerator roadmaps.
U.S. semiconductor market sees a boost; European AI chip makers could feel competitive pressure.
Highlights continued U.S. leadership in advanced AI infrastructure, relevant for global tech investors.
Counterpoint
If adoption of the Groq 3 LPX is slower than expected, the hype could fade and the stock may not see sustained upside.
Key entities
- CompanyNVIDIA
U.S. semiconductor and AI hardware leader.
- CompanyNebius
AI cloud provider planning to use the Groq 3 LPX.



