OpenAI’s upcoming Jalapeño chip looks like it’ll be an inference beast

← Back to the feed

OpenAI’s upcoming Jalapeño chip looks like it’ll be an inference beast

Developing story first seen 3 hours ago

The Register · 3 hours ago

OpenAI has published its most detailed specifications yet for its Jalapeño inference chip, alongside new benchmark figures shared with press ahead of its Hot Chips conference presentation on Tuesday at Stanford. Built with Broadcom, Jalapeño is OpenAI's first custom silicon and is designed purely for running AI models rather than training them, giving it a narrower focus than rival GPUs from Nvidia and AMD. OpenAI says this specialisation delivers higher throughput, lower latency and reduced power draw than current-generation rack systems, though the firm will keep relying on Nvidia and AMD, which remain both suppliers and investors, for training and near-term deployment.

Each Jalapeño chip reportedly delivers 13.4 petaFLOPS at MXFP4 precision with 216GB of HBM4 memory and 15.4 TB/s of bandwidth, scaling to a 128-chip rack offering 1.7 exaFLOPS of 4-bit compute, 27.5TB of HBM4 and nearly 2 petabytes per second of bandwidth. Using SemiAnalysis' InferenceX benchmark suite, OpenAI claims 1.5x to 1.9x greater peak throughput, 1.7x to 3.6x lower latency, and up to 4.1x faster ultra-low-latency inference than comparable Nvidia and AMD systems, though these figures exclude speculative decoding and compare against Nvidia's 2024/2025-era GB200 and GB300 NVL72 racks. OpenAI has not disclosed power consumption, but The Register estimates the chip could use 40-60% less energy than rival GPU systems; Jalapeño is expected to begin shipping later this year, reaching volume production in 2027.

  • OpenAI reveals per-chip and rack-level specs for its Jalapeño AI chip
  • Claims up to 4.1x faster ultra-low-latency inference than Nvidia/AMD rivals
  • Chip ships late 2026, volume production expected in 2027

New here? Start with this

Political and technological unrest at OpenAI, the artificial intelligence company behind ChatGPT, has extended into hardware. Alongside the software it is best known for, OpenAI has been developing its own computer chip, codenamed Jalapeño, built specifically to run AI models once they have already been trained rather than to train them from scratch. Until now, OpenAI has relied on chips made by outside firms such as Nvidia and AMD, which dominate the market for AI computing hardware.

Jalapeño is being made in partnership with Broadcom, a large US chipmaker with experience designing custom silicon for major tech companies. Running an AI model day-to-day, known as inference, is a different job from training it, and a chip built only for that task can potentially be faster and more efficient than general-purpose ones. This matters because running AI systems at scale is hugely expensive and power-hungry, so any gains in speed or efficiency have significant financial and practical consequences for OpenAI and its rivals.

The chip's development sits within a broader race among AI companies to control more of their own technology stack, reducing reliance on external suppliers while those same suppliers, including Nvidia, remain important business partners and investors in OpenAI.

More coverage

AI Technology

Read the full article at the source →