OpenAI's 'Jalapeño' chip beat Nvidia on efficiency: what it actually means
OpenAI showed off its first in-house AI chip and said it does more work per watt than Nvidia's Blackwell. The claim is real, but the fine print matters: these are OpenAI's own benchmarks, against Nvidia's older memory, from a chip that barely ships yet.

At the Hot Chips conference on 25 August, OpenAI detailed Jalapeño, its first in-house AI chip (co-designed with Broadcom and first revealed in June), and published tests showing it does 1.5 to 1.9 times more work per watt than Nvidia's Blackwell chips. That is a genuine result, but three things temper it: the benchmarks are OpenAI's own, not an independent lab's; Jalapeño uses newer memory (HBM4) than the Nvidia parts it beat (HBM3e), so the fair rival is Nvidia's not-yet-widespread Rubin; and Jalapeño is inference-only, barely in production, and for OpenAI's own use. OpenAI also said plainly that it will keep buying a lot of Nvidia.
For years the question hanging over the AI boom has been whether anyone could loosen Nvidia's grip on the chips that run it. On 25 August 2026, at the Hot Chips conference, OpenAI put numbers to its answer: Jalapeño, its first custom AI accelerator, revealed back in June, co-designed with Broadcom and reportedly built on TSMC's 3-nanometre process. Alongside it, OpenAI published benchmarks claiming its chip does meaningfully more AI work per watt than Nvidia's top Blackwell parts. The headline wrote itself, and the coverage duly said OpenAI's chip "beat" Nvidia. The real story is more interesting, and more qualified.
Update, September 2026
Since this published, the timeline moved faster than the piece implied. On Broadcom's 2 September earnings call, its chief executive said Jalapeño had gone from samples into shipment during the quarter ("we also shipped Jalapeño, OpenAI's first-generation custom accelerator"), and set out a hard capacity roadmap: roughly 1.3 gigawatts of Jalapeño deployment planned for 2027, rising to more than 5 gigawatts across Jalapeño and its successor generation in 2028. Broadcom, which now counts six custom-silicon customers (OpenAI, Google, Anthropic and Meta among the named ones), reported $16.7 billion in AI-chip revenue for the quarter and guided sharply higher into 2027 and 2028, a measure of how much of the AI build-out is shifting to made-to-order chips. Two caveats from this piece still stand, though: Nvidia's Rubin, the genuinely like-for-like rival, reached full production and began shipping to customers over the second half of 2026, and there is still no independent, physical head-to-head of Jalapeño against it. The efficiency claim remains OpenAI's own, now backed by a shipping roadmap rather than a benchmark slide alone.
What OpenAI actually claimed
Jalapeño is an inference chip. It runs already-trained models rather than training new ones, which is the cheaper, higher-volume half of the AI workload. In OpenAI's tests, across three open-weight models, it delivered 1.5 to 1.9 times more throughput per watt and 1.7 to 3.6 times lower latency than Nvidia's GB200 and GB300 systems. Richard Ho, who leads OpenAI's hardware effort, presented figures such as a 1.9x efficiency edge on one model and a more than threefold latency improvement on another. Impressively, OpenAI says the team went from first design to a finished chip layout in around nine months.
Those are strong numbers. They are also, and this is the part most headlines skated over, OpenAI's own numbers.
The caveat that changes the story
The benchmarks were run by OpenAI, using a public power-normalised test suite. Analysts at SemiAnalysis watched the runs in person and vouched for what they saw, but they did not run the full suite themselves, and they were blunt about the comparison. Their word for it was "somewhat incomplete and unfair."
The reason is memory. Jalapeño uses HBM4, the newest generation of high-bandwidth memory. The Nvidia chips it was measured against, the GB200 and GB300, use the older HBM3e. Comparing a new-memory chip to old-memory ones flatters the newcomer. The genuinely like-for-like rival is Nvidia's next chip, Rubin, which also uses HBM4, and which is already starting to ship to customers, while Jalapeño was, at that point, still at the engineering-sample stage. There is also a power gap worth naming: Jalapeño is rated at 700 watts and reportedly drew around 550 in testing, while the Nvidia parts it beat draw 1,200 to 1,400 watts. A per-watt win over chips pulling roughly twice the power is real, but it is not the same as being twice as good.
Not a Nvidia replacement, and OpenAI said so
The framing of "OpenAI dumps Nvidia" is wrong on the facts. Jalapeño only does inference, so Nvidia stays essential for the training that creates OpenAI's models in the first place. And OpenAI was explicit: it will "continue to widely deploy accelerators from Nvidia and other partners." Ho put it more plainly still, saying the company continues to need a lot of Nvidia. Jalapeño is a way for OpenAI to shave the cost of running its own services at enormous scale, not a product it will sell or rent to anyone else.
The timeline underlines the point. The first Jalapeño systems are due in OpenAI's own data centres by the end of 2026, in very small numbers, with meaningful volume only through 2027. At its Hot Chips unveiling, the units were engineering samples. So the chip that "beat Nvidia" at Hot Chips will not be doing much real work for a while yet.
Why it still matters
None of that makes Jalapeño unimportant. It is the clearest sign yet that the biggest AI companies would rather design their own silicon than pay Nvidia's margins forever. OpenAI now joins a crowd doing exactly that: Google has its TPUs, Amazon has Trainium (with more than a million chips deployed), Meta has MTIA and Microsoft has Maia. Custom chips are the fastest-growing slice of the AI hardware market. Nvidia still holds roughly 70 per cent of it, but each new in-house design chips away, slowly, at the idea that it is the only option.
The honest read is somewhere between the hype and the dismissal. OpenAI built a genuinely efficient inference chip, faster to design than anyone expected, and good enough that a sceptical analyst confirmed what they saw. It did not, at Hot Chips, beat Nvidia in a fair fight, because the fair fight is against Rubin; analysts at SemiAnalysis modelled that match-up and had Jalapeño edging Rubin on efficiency but roughly even on cost per token, and no independent, physical head-to-head has happened yet. What Jalapeño really shows is that the economics of the AI boom are pushing its biggest spenders to build their way out from under their most important supplier. Whether that works is a 2027 question.





