OpenAI’s upcoming Jalapeño chip looks like it’ll be an inference beast
Developing story first seen 3 hours ago
OpenAI has published its most detailed specifications yet for its Jalapeño inference chip, alongside new benchmark figures shared with press ahead of its Hot Chips conference presentation on Tuesday at Stanford. Built with Broadcom, Jalapeño is OpenAI's first custom silicon and is designed purely for running AI models rather than training them, giving it a narrower focus than rival GPUs from Nvidia and AMD. OpenAI says this specialisation delivers higher throughput, lower latency and reduced power draw than current-generation rack systems, though the firm will keep relying on Nvidia and AMD, which remain both suppliers and investors, for training and near-term deployment.
Each Jalapeño chip reportedly delivers 13.4 petaFLOPS at MXFP4 precision with 216GB of HBM4 memory and 15.4 TB/s of bandwidth, scaling to a 128-chip rack offering 1.7 exaFLOPS of 4-bit compute, 27.5TB of HBM4 and nearly 2 petabytes per second of bandwidth. Using SemiAnalysis' InferenceX benchmark suite, OpenAI claims 1.5x to 1.9x greater peak throughput, 1.7x to 3.6x lower latency, and up to 4.1x faster ultra-low-latency inference than comparable Nvidia and AMD systems, though these figures exclude speculative decoding and compare against Nvidia's 2024/2025-era GB200 and GB300 NVL72 racks. OpenAI has not disclosed power consumption, but The Register estimates the chip could use 40-60% less energy than rival GPU systems; Jalapeño is expected to begin shipping later this year, reaching volume production in 2027.
- OpenAI reveals per-chip and rack-level specs for its Jalapeño AI chip
- Claims up to 4.1x faster ultra-low-latency inference than Nvidia/AMD rivals
- Chip ships late 2026, volume production expected in 2027
New here? Start with this
Political and technological unrest at OpenAI, the artificial intelligence company behind ChatGPT, has extended into hardware. Alongside the software it is best known for, OpenAI has been developing its own computer chip, codenamed Jalapeño, built specifically to run AI models once they have already been trained rather than to train them from scratch. Until now, OpenAI has relied on chips made by outside firms such as Nvidia and AMD, which dominate the market for AI computing hardware.
Jalapeño is being made in partnership with Broadcom, a large US chipmaker with experience designing custom silicon for major tech companies. Running an AI model day-to-day, known as inference, is a different job from training it, and a chip built only for that task can potentially be faster and more efficient than general-purpose ones. This matters because running AI systems at scale is hugely expensive and power-hungry, so any gains in speed or efficiency have significant financial and practical consequences for OpenAI and its rivals.
The chip's development sits within a broader race among AI companies to control more of their own technology stack, reducing reliance on external suppliers while those same suppliers, including Nvidia, remain important business partners and investors in OpenAI.
Full account
OpenAI has given the clearest look yet at its first custom AI chip, codenamed Jalapeño, unveiling fresh performance details at the Hot Chips semiconductor conference at Stanford University and in a company blog post published on Tuesday. Developed jointly with Broadcom, the chip is an application-specific integrated circuit built purely for AI inference — that is, running an already-trained model to generate responses or power an agent — rather than for the training work that still dominates OpenAI's use of Nvidia and AMD hardware. Company hardware vice-president Richard Ho told reporters the design was intended to sidestep the usual trade-off between speed and volume, offering what he called the 'best of both worlds' of low latency and high throughput.
OpenAI shared benchmark figures drawn from InferenceX, a third-party inference testing suite, comparing systems built on Jalapeño against the strongest results currently logged for Nvidia's GB200 and GB300 chips. Across three models — GPT-OSS-120B, DeepSeek R1 and Kimi K2.5 — OpenAI said its chip completed between 1.5 and 1.9 times more useful work at peak throughput while cutting overall response times by a factor of 1.7 to 3.6. For the fastest-turnaround workloads, sometimes described as ultra-low-latency inference and increasingly seen as a priority for AI infrastructure providers, OpenAI put the improvement as high as two to four times. At the level of a full server rack, a Jalapeño system linking 128 of the chips is said to offer roughly 1.7 exaFLOPS of low-precision compute, around 27.5 terabytes of high-bandwidth memory, and just under two petabytes per second of memory bandwidth — a spec sheet that, notably, prioritises memory bandwidth over raw compute, reflecting how inference workloads are typically bottlenecked by data movement rather than calculation.
Despite the strong early figures, OpenAI has been careful to frame Jalapeño as a complement to, rather than a replacement for, its existing hardware relationships. Ho reiterated that the company's broader compute strategy still relies on 'very good partners' such as Nvidia, whose chips remain better suited to the more varied demands of model training and offer greater general-purpose flexibility. The first Jalapeño units are expected to ship in small quantities before the end of this year, with production ramping up through 2027 — putting it on a similar timeline to Nvidia's next-generation Rubin platform and AMD's MI455X. OpenAI has already signalled it intends to keep iterating, with second- and third-generation versions of the chip in development.
Commentators noted that the exercise doubles as a signal of OpenAI's growing appetite for hardware independence, following its earlier moves to diversify beyond a single chip supplier. Because Jalapeño is designed solely for inference rather than as a do-everything accelerator, it can in principle be leaner and more efficient at that one task than general-purpose GPUs, without sacrificing the programmability needed to run a range of different models — a contrast drawn with more narrowly optimised, model-specific rivals. OpenAI has not yet disclosed how much power a full Jalapeño rack consumes, an omission that leaves open questions about how the efficiency gains translate into running costs at scale.
Where outlets differ
The two accounts emphasise different figures from the same InferenceX results: one focuses on peak throughput and end-to-end latency multiples (1.5–1.9x work, 1.7–3.6x lower latency), while the other frames the comparable numbers as 'AI work per watt', suggesting some ambiguity or differing interpretation of exactly what OpenAI's benchmark claims measure.
One source gives detailed system-level specifications (128 chips per rack, 1.7 exaFLOPS, 27.5TB of HBM4, near 2PB/s bandwidth) and compares these directly against AMD and Nvidia rack systems, including Nvidia's Helios; the other source omits these hardware specifics entirely, focusing instead on the briefing, quotes from OpenAI's hardware VP Richard Ho, and deployment timing.
One account is more sceptical/analytical in tone, describing the InferenceX test as 'presumably unofficial' and speculating about power consumption, whereas the other reports OpenAI's statements more straightforwardly as a company announcement, drawing on an official blog post and press briefing.
One source situates the chip within a competitive timeline against Nvidia's Rubin and AMD's MI455X GPUs, both expected in early 2027, while the other does not directly compare Jalapeño's ramp to rival product roadmaps.
Only one source notes that Jalapeño was first teased/introduced in June and confirms OpenAI's plan to develop second- and third-generation chips.
More coverage