OpenAI's Jalapeño

A deep dive into the first custom OpenAI inference ASIC

Summary

On August 25, 2026, at the Hot Chips conference at Stanford, OpenAI publicly detailed Jalapeño, its first custom-built AI inference accelerator, co-developed with Broadcom and integrated by Celestica 123. The chip is the opening salvo in a multi-generation roadmap that re-architects inference silicon around three-phase agentic workloads (prefill, draft, verification) instead of the classic two-phase prefill/decode model that Nvidia GPUs target 12. In public benchmarks against Nvidia GB200 and GB300 using SemiAnalysis's InferenceX suite, Jalapeño reaches 1.5×–1.9× higher throughput per watt and 1.7×–3.6× lower end-to-end latency on GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T 123. Beyond raw specs, the story is the methodology: more than half of the core was authored in the open-source XLS hardware language and optimized by OpenAI models themselves, taking the project from initial RTL to tape-out in approximately nine months 12.

Key Takeaways

  • A balanced inference platform, not a GPU replacement. Jalapeño is a 700 W package with 216 GiB HBM4, ~15.4 TB/s HBM bandwidth, 13.4 PFLOPS MXFP4, and 3.4 PFLOPS MXFP8, scaled to 27 EFLOPS and 432 TiB at the 2,048-ASIC pod level 1.
  • NUMA-style, sliced architecture. 64 core slices, each paired with its own HBM slice, connected by a low-latency collective network for register-to-register operand movement plus a deliberately "anemic" general-purpose NoC 12.
  • Three-phase agentic inference. OpenAI splits inference into compute-bound prefill, latency-bound draft (small speculative model), and bandwidth-bound spec-verify decode, then optimizes one chip for all three instead of building heterogeneous fleets 12.
  • AI-designed AI silicon. OpenAI used its own models and XLS to write the bulk of the RTL, then had those models sweep power-performance-area (PPA), producing reported gains such as a 56% improvement on a BF16 multiplier and a 10% matrix-unit area reduction vs. human baseline 12.
  • AI-programmed AI silicon. Software runs on top of the open-source Triton/Gluon stack, with OpenAI's models then auto-tuning kernels, yielding 1.5×–1.8× speedups over expert-written attention and MoE kernels on chip 2.
  • Competitive, but not a dethroning. SemiAnalysis notes that comparing Jalapeño against GB200/GB300 favors OpenAI; the more forward-looking comparison is Nvidia's Rubin generation 3.
  • Gen 2 and Gen 3 are already in flight. OpenAI told Hot Chips attendees that Jalapeño is "Gen 1," Gen 2 is in development heading toward tape-out, and Gen 3 is "operational" (planned) 1.

Main Content

1. Why a Custom Inference Chip, Why Now

Jalapeño was formally introduced on June 24, 2026 as a Broadcom-co-developed, Celestica-integrated inference ASIC built "from scratch around modern LLM inference" rather than adapted from a broader compute IP 3. Three motivations stand out from OpenAI's framing at Hot Chips:

  1. Inference is the recurring cost center for AI products. Training is a one-shot investment, while serving a frontier model to hundreds of millions of users is an ongoing power, networking, and silicon bill 3.
  2. Workloads are increasingly agentic. A single agentic request spans prefill, speculative drafting, and verification, hitting different bottlenecks in sequence, which makes a single balanced chip more efficient than specialized fleets where some accelerators sit idle 12.
  3. Hardware, software, and model co-design is now tractable. OpenAI argues that only the operator of both the models and the silicon can fully exploit PPA, kernel, and scheduling opportunities 12.

This places Jalapeño in the same strategic category as Google's TPU, AWS's Trainium/Inferentia, and Microsoft's Maia: a vertical integration move to control inference economics and supply diversity 3.

2. Architecture Deep Dive

2.1 The Package

  • Package power: 700 W TDP, currently clocked at 1.70 GHz in OpenAI's labs with a roadmap to 1.80 GHz 1.
  • Compute: Up to 3.4 PFLOPS MXFP8 and up to 13.4 PFLOPS MXFP4 12.
  • Memory: 216 GiB HBM4 at ~15.4 TB/s bandwidth 12.

On paper these numbers trail Nvidia Blackwell Ultra's 10 FP8 PFLOPS and 20/15 NVFP4 PFLOPS, but OpenAI's central claim is that raw peak is not the bottleneck for inference 1.

2.2 The Sliced, NUMA-Style Core

Jalapeño's compute die is partitioned into 64 core slices, each paired with a slice of HBM. This gives every slice a fast, predictable local memory view and removes the unified-memory contention that bottlenecks large GPU packages 12.

Two on-chip fabrics tie the slices together:

  • A specialized collective network that moves operands "register-to-register, with zero conflicts" for common cross-core patterns like attention and MoE all-to-all 12.
  • A general-purpose NoC that OpenAI explicitly describes as more "anemic" than a typical design, used only for remote/global traffic so performance-critical paths never fight for it 12.

2.3 The Scale-Up Fabric

  • Local domain: 128 ASICs at 600 GB/s interconnect 12.
  • Global domain: 2,048 ASICs across 16 racks via a "half-flattened" two-level Clos built on Broadcom Tomahawk 6 Ethernet switches at 200 GB/s per ASIC 12.
  • Aggregate system: 27 EFLOPS MXFP4, 432 TiB HBM4, and 32 PB/s aggregate memory bandwidth 12.

Higher bandwidth is reserved for tensor-parallel traffic, lower bandwidth for expert-parallel communication, and both are tuned for low latency 12.

3. Three-Phase Agentic Inference

Nvidia's standard decomposition of an LLM request is prefill (compute-bound) + decode (memory-bound), which is why Nvidia segregates accelerators accordingly (e.g., Rubin CPX for prefill with GDDR7) 1.

OpenAI argues inference is really three phases:

  1. Prefill (compute-bound)
  2. Draft (latency-bound): a small speculative model runs at ultra-low batch.
  3. Spec-verify decode (bandwidth-bound, with bursty MoE traffic).

Because the mix of these phases shifts with model, context length, token efficiency, and software stack, OpenAI rejects specialized heterogeneous fleets and ships one balanced ASIC where idle units are power-gated between phases 12. KV cache stays local to the chip, eliminating the cross-network KV movement that bogs down specialized accelerators at phase boundaries 12.

4. Performance: Pareto Frontier, Not Peak FLOPS

OpenAI benchmarked using SemiAnalysis's InferenceX, normalizing results to package TDP (Jalapeño 700 W vs. GB200 1,200 W, GB300 and MI355X 1,400 W) and comparing across the full latency-vs-throughput curve 12.

Reported wins (Jalapeño vs. GB200/GB300):

Workload Peak throughput / watt End-to-end latency
GPT-OSS 120B ~1.9× ~1.7× lower
DeepSeek R1 670B ~1.7× ~3.6× lower
Kimi K2.5 1T ~1.5× ~3.4× lower

Source: 12.

OpenAI also reports sub-millisecond token-to-token latency on frontier models at economical throughput, with multi-token prediction projected to add another 3×–5× latency improvement at iso-efficiency 2.

Caveat from SemiAnalysis: the comparison is mostly against Blackwell (GB200/GB300), not the newer Rubin generation. SemiAnalysis calls Jalapeño a credible first-gen custom accelerator, not a dethroning of Nvidia 3.

5. AI Designing AI, and AI Programming AI

5.1 Design

OpenAI's team built most of the compute die from scratch in XLS (Google's open-source hardware-description and high-level synthesis toolchain) plus Verilog; only some interface IP was reused from Broadcom 12.

Reported PPA gains from AI-assisted sweeps over human baselines:

  • +56% on a BF16 multiplier
  • +21% on an FP4 dot-product block
  • +10% on an FP32 accumulator
  • −10% matrix-unit area
  • −8% SIMD-unit area 12

5.2 Timeline

  • Feb 2025: initial RTL work begins
  • Nov 2025: tape-out
  • May 2026: first silicon in OpenAI's labs; Codex running on Jalapeño the same month 12

OpenAI frames this as a roughly 9-month design-to-tape-out, while SemiAnalysis counts a longer ~16 months from early team formation 13.

5.3 Software and Kernels

Jalapeño is programmed through Triton / Gluon, exposing physical placement explicitly so AI search can handle mapping, scheduling, and pipelining that would be onerous for humans 12. OpenAI reports its internal AI auto-tuner produces attention and MoE kernels that run 1.5×–1.8× faster than expert-written implementations, validated end-to-end on-chip 2.

6. Roadmap and Supply Strategy

  • Gen 1 (Jalapeño): OpenAI reportedly plans to begin deploying it in its own infrastructure by the end of 2026 4.
  • Gen 2: in development, heading toward tape-out, targeting better performance per watt 12.
  • Gen 3: already "operational" (or "planned," per slide) according to Richard Ho at Hot Chips, focused on economical low-latency serving 12.

Crucially, OpenAI has stated explicitly that it intends to continue deploying Nvidia and other partners' accelerators alongside Jalapeño for both training and inference; the goal is supply diversity and inference economics, not a unilateral Nvidia exit 3.

7. What This Means for AI Engineering

For builders and operators, the practical implications of the Jalapeño disclosure are:

  • Pareto-frontier thinking wins. OpenAI's benchmark methodology (latency-vs-throughput at normalized TDP) is becoming a more useful comparison axis than peak FLOPS or theoretical HBM bandwidth 12.
  • Sliced, NUMA-style memory is the new default. Pairing compute slices with local HBM slices, plus a dedicated low-latency collective fabric, is likely to be replicated across competitors in 2027.
  • AI-designed silicon is no longer a demo. More than half the core authored in XLS and tuned by OpenAI's own models is now first-gen silicon already running Codex in OpenAI's labs, with kernel-level autotuners in the same boat 12.
  • Agentic inference shapes silicon. Expect next-gen accelerators to budget silicon for small speculative models and verification, not just prefill and decode.

Conclusion

OpenAI's Jalapeño is a credible, full-stack custom inference ASIC from a frontier AI lab. It is not a GPU, and it is not a Nvidia killer; it is a deliberately balanced, sliced, AI-designed inference platform tuned for the three-phase reality of agentic requests, scaled across 2,048-ASIC pods. The deeper story is the feedback loop: OpenAI's models designed the chip, OpenAI's models now program the kernels on the chip, and the resulting economics will make the next generation of OpenAI's models cheaper and faster to serve. For AI engineering teams, the takeaway is to start benchmarking and architecting inference along Pareto frontiers (latency vs. tokens per joule) rather than peak FLOPS, and to plan for hardware that treats prefill, draft, and verify as first-class citizens.

Sources

  1. Tom's Hardware, "Hot Chips 2026: OpenAI's Jalapeño AI ASIC unpacked" (Aug 27, 2026) — https://www.tomshardware.com/tech-industry/artificial-intelligence/hot-chips-2026-openais-jalapeno-ai-asic-unpacked-accelerator-developed-using-ai-achieves-efficiency-and-throughput-gains-against-power-hungry-blackwell
  2. ServeTheHome, "OpenAI Jalapeno Custom AI ASIC at Hot Chips 2026" (Aug 26, 2026) — https://www.servethehome.com/openai-jalapeno-asic-at-hot-chips-2026/
  3. Quantilus, "The Custom AI Chip That Could Reshape the Inference Race" (Aug 26, 2026) — https://quantilus.com/article/the-custom-ai-chip-that-could-reshape-the-inference-race/
  4. NDTV Profit, "'We Made A Chip And It Is Fast,' Says Sam Altman As OpenAI Unveils Custom Inference Chip 'Jalapeno'" (Aug 26, 2026) — https://www.ndtvprofit.com/technology/we-made-a-chip-and-it-is-fast-says-sam-altman-as-openai-unveils-custom-inference-chip-jalapeno-11959287

Keep reading

All posts →

Be first in when doors open.

Get early access