The one that matters
Nvidia's dedicated inference chip is now in full production
Nvidia says Groq 3 LPX, its inference-only accelerator, has entered full production, posting 3,431 tokens/sec on Gemma 4 31B with a 100,000-token input - which it claims is 4x the next public endpoint. Nebius signed as first customer, and SpaceX will fly a space-optimised Vera Rubin NVL72 to orbit next year.
Why it matters
Agent workloads are dominated by inference, not training - long contexts, many calls, tight latency. Silicon built only for that, shipping at volume, reprices every serious agent deployment. The Register's read of the same benchmark is a useful corrective: one spectacular number on one model is not a price list.
