OpenAI's Jalapeño Chip Beats Nvidia's Best in First Official Benchmarks: The Full Numbers
On Tuesday, August 25, 2026, OpenAI published the first official performance results for its custom "Jalapeño" chip, claiming "industry-leading speed and efficiency in AI inference" — beating the best recorded results from Nvidia's GB200 and GB300 systems on SemiAnalysis' InferenceX benchmarking platform. The numbers are concrete: 1.5 to 1.9 times more AI work per watt, and 1.7 to 3.6 times lower end-to-end latency, across three major open-weight models. This is the first published benchmark for the chip OpenAI unveiled in June, and the first hard public evidence that the company's custom-silicon strategy is starting to pay off.
This piece breaks down the numbers exactly as published, how the test was run, what it practically means for ChatGPT and agent users worldwide, and — just as importantly — where the limits of these results are.
The Announced Numbers: What Jalapeño's First Test Shows
According to OpenAI's official post, the chip was evaluated on InferenceX — the inference benchmarking platform run by SemiAnalysis — and compared against the best results recorded on the platform at the time of testing, all of which were achieved on Nvidia GB200 or GB300 systems. The test covered three widely used open-weight models:
- GPT-OSS 120B: OpenAI's mid-size open-weight model.
- DeepSeek R1: the well-known Chinese reasoning model.
- Kimi K2.5 1T: the newest trillion-parameter release from Moonshot AI.
Across these three models, OpenAI says Jalapeño delivered 1.5 to 1.9 times more AI work per watt than the comparison systems — meaning 50% to 90% more inference for every unit of electricity consumed. It also recorded 1.7 to 3.6 times lower end-to-end latency, meaning the total time from request to completed response shrank dramatically.

Source: OpenAI — Jalapeño's first results
"Best of Both Worlds": What OpenAI Said
In a briefing with reporters after the announcement, OpenAI hardware vice president Richard Ho said Jalapeño offers the "best of both worlds": lower latency and higher throughput at the same time, whereas AI systems typically "have to make a trade-off between the two." The practical meaning, according to Ho, is "faster responses, more responsive agents, and more reliable access as the demand grows" — a direct nod to the exploding inference load that AI agents are creating.
The core point: the historical trade-off between latency (how quickly the next token appears) and throughput (how many users the system serves at once) is exactly what Jalapeño was designed to break, and the numbers above suggest it is at least close to doing so in the published test.

Source: OpenAI — Jalapeño's first results
What Exactly Is Jalapeño?
Jalapeño is an ASIC — an application-specific integrated circuit built for exactly one job — developed by OpenAI in partnership with Broadcom. That job is inference: running trained models to answer users and execute tasks, not training them. The chip was first introduced in June 2026 as part of OpenAI's plan to reduce its dependence on Nvidia and rein in the enormous compute costs of running its services.
The wider context matters. The AI silicon race has intensified sharply this year: Meta is preparing to manufacture its new Iris chip in September, as we covered in our earlier reporting, Google keeps expanding its TPU fleet, and Amazon is developing Trainium. OpenAI entering the arena with genuinely competitive performance numbers signals that the AI inference market is shifting from a near-monopoly toward multi-player competition — which, historically, is what drives prices down and quality up.
What This Means for You as a User
A chip war may feel far removed from a ChatGPT user in Riyadh, Dubai, or anywhere else, but its effects arrive directly:
- Faster peak-hour responses: if Jalapeño serves part of the inference load, the first beneficiary is response time during congestion — including Gulf evening hours, when global demand typically strains capacity.
- More reliable agents: AI agents chain thousands of sequential inference steps, and any per-step slowdown compounds into the final result. Cutting latency 1.7 to 3.6 times means agents that complete longer tasks without stalling.
- Energy efficiency = sustainable scaling: inference datacenters serving the region are growing fast, and every watt saved translates into capacity to serve more users without outages — what Ho called "more reliable access as demand grows."
- Do not expect immediate price cuts: OpenAI announced no pricing change tied to the chip, and we will not predict one. Any pricing effect — if it comes — would materialize as volumes scale in 2027.
Quick Comparison: Jalapeño vs Nvidia Systems in the Published Test
| Metric (as announced by OpenAI) | Jalapeño | Comparison systems (GB200/GB300) |
|---|---|---|
| Work per watt (across 3 models) | 1.5–1.9x | Recorded baseline |
| End-to-end latency | 1.7–3.6x lower | Recorded baseline |
| Models tested | GPT-OSS 120B, DeepSeek R1, Kimi K2.5 1T | Same models |
| Party publishing the numbers | OpenAI itself, on the SemiAnalysis platform | — |
Note that the comparison is not against a specific Nvidia system OpenAI bought and ran itself, but against "the best results recorded" on InferenceX at test time — a precise phrasing worth keeping in mind when reading enthusiastic headlines elsewhere.
Honest Limitations: What the Test Does Not Show
- The numbers come from OpenAI itself: the test ran on an independent platform (InferenceX), but the party publishing and interpreting the results is their primary beneficiary. Full independence would require third-party re-testing, which has not happened yet.
- Initial volumes are small: OpenAI will deploy Jalapeño in "small volumes" by the end of 2026 and ramp up into 2027, without disclosing planned chip counts. Any real effect on service performance will accumulate gradually.
- Nvidia is not going away: Ho confirmed OpenAI's compute strategy includes "very good partners" like Nvidia, and the company does not intend to replace its entire lineup with Jalapeño.
- Generations two and three are in development: what we saw is the first generation, with teams already building the next two — meaning the long-term competition has barely started.
The Deployment Timeline
Per Richard Ho: limited deployment "by the end of this year" — before the close of 2026 — followed by a volume ramp "into 2027." Alongside the benchmark results, OpenAI published a broader piece on its vision for "abundant intelligence" and its full compute stack, a signal that the chip is one pillar of a deliberate strategy rather than an experiment.
Frequently Asked Questions About OpenAI's Jalapeño Chip
Did Jalapeño really beat Nvidia?
In the test published August 25, 2026, Jalapeño recorded higher results than the best recorded GB200/GB300 results on the InferenceX platform across three models. Those are OpenAI's own published numbers and have not yet been independently confirmed.
When will I notice the chip as a user?
Not immediately; deployment starts in small volumes before the end of 2026 and scales through 2027. The expected effect is gradual: faster responses and more stable agents during peak hours.
Who manufactures the Jalapeño chip?
OpenAI developed it with Broadcom as an inference-focused ASIC — for running trained models, not training them.
Is OpenAI dropping Nvidia?
No. OpenAI's hardware VP said Nvidia remains a "very good partner" and Jalapeño will not replace the company's entire chip lineup, with two more generations in development.
Does the chip affect ChatGPT pricing?
There is no official announcement of any price change tied to the chip. Higher energy efficiency opens the door to future cost improvements, but any prediction today would be speculation.
To compare AI tools you can actually use today, browse Truescho's AI tools directory.