The one that matters
OpenAI's first chip posts numbers Nvidia has to answer
OpenAI says Jalapeño, the inference ASIC it co-designed with Broadcom in 16 months, delivered 1.5-1.9x more AI work per watt and 1.7-3.6x lower latency than Nvidia hardware across GPT-OSS, DeepSeek R1, and Kimi K2.5 1T. SemiAnalysis' independent breakdown broadly backs the efficiency story.
Why it matters
Custom inference silicon from the biggest model shop reprices the agent economy - though 30-year chip veterans are openly sceptical a v1 part holds these numbers at production scale. Either way, Nvidia's per-watt moat now has a named challenger with published benchmarks.
