OpenAI publishes Jalapeño inference benches against Nvidia GB200 and GB300
OpenAI is publishing its own inference numbers against last-gen Nvidia silicon.
The Verge reports that OpenAI says Jalapeño completes tasks more efficiently and returns responses faster than other AI systems, according to a Tuesday blog post. Richard Ho, OpenAI’s hardware vice president, told Verge the chip offers the “best of both worlds” on latency and throughput, because AI systems typically have to trade one for the other.
Jalapeño is an ASIC made with Broadcom. Verge: it is designed for inference, running a trained model to complete a task or deploy an agent. The Next Web: it cannot train models at all.
The benches used InferenceX. Verge: the test compared Jalapeño against the best results recorded at the time, which were with Nvidia’s GB200 or GB300 superchips.
OpenAI says Jalapeño delivered 1.5 to 1.9 times more AI work per watt across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T than those comparison systems.
The same Verge post says 1.7 to 3.6 times lower end-to-end latency across the three models.
The Next Web: these are still OpenAI’s numbers.
SemiAnalysis went to the lab. They say all numbers were provided by OpenAI.
They verified some InferenceX runs in person. They did not run the full suite, and they have not seen AgentX.
SemiAnalysis argues a Blackwell comparison is incomplete because Vera Rubin uses HBM4 and is shipping to customers now, while Jalapeño is still engineering samples scheduled to ramp through 2027.
The Next Web: Jalapeño was not tested against Rubin, which has just started shipping.
The Next Web and SemiAnalysis put Jalapeño at 700 watts, a figure Verge does not print.
SemiAnalysis: design work began in the middle of 2024, from initial team hiring to manufacturing tape-out in about 16 months.
OpenAI plans to deploy Jalapeño in small volumes by the end of this year, then ramp into 2027. The company does not say how many chips it plans to deploy next year.
Ho told Verge OpenAI does not expect to replace its entire chip lineup. The compute strategy still includes “very good partners,” like Nvidia. He told The Next Web: “Nvidia is a really good partner, and we continue to need a lot of Nvidia.”
OpenAI will keep developing the second and third generations.
Jalapeño was not tested against Rubin. SemiAnalysis has not seen the full suite.
The foil is last year’s rack.

Pingback: Instagram's opt-in, Malone's empty chair, and Huawei's Egypt bid