# OpenAI's Jalapeño

A deep dive into the first custom OpenAI inference ASIC

Aug 31, 2026 · Inference · 8 min read · https://arclyx.ai/blog/openai-jalapeno

## Summary

On August 25, 2026, at the Hot Chips conference at Stanford, OpenAI publicly detailed Jalapeño, its first custom-built AI inference accelerator, co-developed with Broadcom and integrated by Celestica [Source 1][Source 2][Source 3]. The chip is the opening salvo in a multi-generation roadmap that re-architects inference silicon around three-phase agentic workloads (prefill, draft, verification) instead of the classic two-phase prefill/decode model that Nvidia GPUs target [Source 1][Source 2]. In public benchmarks against Nvidia GB200 and GB300 using SemiAnalysis's InferenceX suite, Jalapeño reaches 1.5×–1.9× higher throughput per watt and 1.7×–3.6× lower end-to-end latency on GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T [Source 1][Source 2][Source 3]. Beyond raw specs, the story is the methodology: more than half of the core was authored in the open-source XLS hardware language and optimized by OpenAI models themselves, taking the project from initial RTL to tape-out in approximately nine months [Source 1][Source 2].

## Key Takeaways

- **A balanced inference platform, not a GPU replacement.** Jalapeño is a 700 W package with 216 GiB HBM4, ~15.4 TB/s HBM bandwidth, 13.4 PFLOPS MXFP4, and 3.4 PFLOPS MXFP8, scaled to 27 EFLOPS and 432 TiB at the 2,048-ASIC pod level [Source 1].
- **NUMA-style, sliced architecture.** 64 core slices, each paired with its own HBM slice, connected by a low-latency collective network for register-to-register operand movement plus a deliberately "anemic" general-purpose NoC [Source 1][Source 2].
- **Three-phase agentic inference.** OpenAI splits inference into compute-bound prefill, latency-bound draft (small speculative model), and bandwidth-bound spec-verify decode, then optimizes one chip for all three instead of building heterogeneous fleets [Source 1][Source 2].
- **AI-designed AI silicon.** OpenAI used its own models and XLS to write the bulk of the RTL, then had those models sweep power-performance-area (PPA), producing reported gains such as a 56% improvement on a BF16 multiplier and a 10% matrix-unit area reduction vs. human baseline [Source 1][Source 2].
- **AI-programmed AI silicon.** Software runs on top of the open-source Triton/Gluon stack, with OpenAI's models then auto-tuning kernels, yielding 1.5×–1.8× speedups over expert-written attention and MoE kernels on chip [Source 2].
- **Competitive, but not a dethroning.** SemiAnalysis notes that comparing Jalapeño against GB200/GB300 favors OpenAI; the more forward-looking comparison is Nvidia's Rubin generation [Source 3].
- **Gen 2 and Gen 3 are already in flight.** OpenAI told Hot Chips attendees that Jalapeño is "Gen 1," Gen 2 is in development heading toward tape-out, and Gen 3 is "operational" (planned) [Source 1].

## Main Content

### 1. Why a Custom Inference Chip, Why Now

Jalapeño was formally introduced on June 24, 2026 as a Broadcom-co-developed, Celestica-integrated inference ASIC built "from scratch around modern LLM inference" rather than adapted from a broader compute IP [Source 3]. Three motivations stand out from OpenAI's framing at Hot Chips:

1. Inference is the recurring cost center for AI products. Training is a one-shot investment, while serving a frontier model to hundreds of millions of users is an ongoing power, networking, and silicon bill [Source 3].
2. Workloads are increasingly agentic. A single agentic request spans prefill, speculative drafting, and verification, hitting different bottlenecks in sequence, which makes a single balanced chip more efficient than specialized fleets where some accelerators sit idle [Source 1][Source 2].
3. Hardware, software, and model co-design is now tractable. OpenAI argues that only the operator of both the models and the silicon can fully exploit PPA, kernel, and scheduling opportunities [Source 1][Source 2].

This places Jalapeño in the same strategic category as Google's TPU, AWS's Trainium/Inferentia, and Microsoft's Maia: a vertical integration move to control inference economics and supply diversity [Source 3].

### 2. Architecture Deep Dive

#### 2.1 The Package

- **Package power:** 700 W TDP, currently clocked at 1.70 GHz in OpenAI's labs with a roadmap to 1.80 GHz [Source 1].
- **Compute:** Up to 3.4 PFLOPS MXFP8 and up to 13.4 PFLOPS MXFP4 [Source 1][Source 2].
- **Memory:** 216 GiB HBM4 at ~15.4 TB/s bandwidth [Source 1][Source 2].

On paper these numbers trail Nvidia Blackwell Ultra's 10 FP8 PFLOPS and 20/15 NVFP4 PFLOPS, but OpenAI's central claim is that raw peak is not the bottleneck for inference [Source 1].

#### 2.2 The Sliced, NUMA-Style Core

Jalapeño's compute die is partitioned into **64 core slices**, each paired with a slice of HBM. This gives every slice a fast, predictable local memory view and removes the unified-memory contention that bottlenecks large GPU packages [Source 1][Source 2].

Two on-chip fabrics tie the slices together:

- **A specialized collective network** that moves operands "register-to-register, with zero conflicts" for common cross-core patterns like attention and MoE all-to-all [Source 1][Source 2].
- **A general-purpose NoC** that OpenAI explicitly describes as more "anemic" than a typical design, used only for remote/global traffic so performance-critical paths never fight for it [Source 1][Source 2].

#### 2.3 The Scale-Up Fabric

- **Local domain:** 128 ASICs at 600 GB/s interconnect [Source 1][Source 2].
- **Global domain:** 2,048 ASICs across 16 racks via a "half-flattened" two-level Clos built on Broadcom Tomahawk 6 Ethernet switches at 200 GB/s per ASIC [Source 1][Source 2].
- **Aggregate system:** 27 EFLOPS MXFP4, 432 TiB HBM4, and 32 PB/s aggregate memory bandwidth [Source 1][Source 2].

Higher bandwidth is reserved for tensor-parallel traffic, lower bandwidth for expert-parallel communication, and both are tuned for low latency [Source 1][Source 2].

### 3. Three-Phase Agentic Inference

Nvidia's standard decomposition of an LLM request is **prefill (compute-bound) + decode (memory-bound)**, which is why Nvidia segregates accelerators accordingly (e.g., Rubin CPX for prefill with GDDR7) [Source 1].

OpenAI argues inference is really three phases:

1. **Prefill** (compute-bound)
2. **Draft** (latency-bound): a small speculative model runs at ultra-low batch.
3. **Spec-verify decode** (bandwidth-bound, with bursty MoE traffic).

Because the mix of these phases shifts with model, context length, token efficiency, and software stack, OpenAI rejects specialized heterogeneous fleets and ships **one balanced ASIC where idle units are power-gated between phases** [Source 1][Source 2]. KV cache stays local to the chip, eliminating the cross-network KV movement that bogs down specialized accelerators at phase boundaries [Source 1][Source 2].

### 4. Performance: Pareto Frontier, Not Peak FLOPS

OpenAI benchmarked using **SemiAnalysis's InferenceX**, normalizing results to package TDP (Jalapeño 700 W vs. GB200 1,200 W, GB300 and MI355X 1,400 W) and comparing across the full latency-vs-throughput curve [Source 1][Source 2].

Reported wins (Jalapeño vs. GB200/GB300):

| Workload | Peak throughput / watt | End-to-end latency |
|---|---|---|
| GPT-OSS 120B | ~1.9× | ~1.7× lower |
| DeepSeek R1 670B | ~1.7× | ~3.6× lower |
| Kimi K2.5 1T | ~1.5× | ~3.4× lower |

Source: [Source 1][Source 2].

OpenAI also reports **sub-millisecond token-to-token latency** on frontier models at economical throughput, with multi-token prediction projected to add another 3×–5× latency improvement at iso-efficiency [Source 2].

**Caveat from SemiAnalysis:** the comparison is mostly against Blackwell (GB200/GB300), not the newer Rubin generation. SemiAnalysis calls Jalapeño a credible first-gen custom accelerator, not a dethroning of Nvidia [Source 3].

### 5. AI Designing AI, and AI Programming AI

#### 5.1 Design

OpenAI's team built most of the compute die from scratch in **XLS (Google's open-source hardware-description and high-level synthesis toolchain)** plus Verilog; only some interface IP was reused from Broadcom [Source 1][Source 2].

Reported PPA gains from AI-assisted sweeps over human baselines:

- **+56%** on a BF16 multiplier
- **+21%** on an FP4 dot-product block
- **+10%** on an FP32 accumulator
- **−10%** matrix-unit area
- **−8%** SIMD-unit area [Source 1][Source 2]

#### 5.2 Timeline

- **Feb 2025:** initial RTL work begins
- **Nov 2025:** tape-out
- **May 2026:** first silicon in OpenAI's labs; **Codex running on Jalapeño** the same month [Source 1][Source 2]

OpenAI frames this as a roughly **9-month** design-to-tape-out, while SemiAnalysis counts a longer ~16 months from early team formation [Source 1][Source 3].

#### 5.3 Software and Kernels

Jalapeño is programmed through **Triton / Gluon**, exposing physical placement explicitly so AI search can handle mapping, scheduling, and pipelining that would be onerous for humans [Source 1][Source 2]. OpenAI reports its internal AI auto-tuner produces attention and MoE kernels that run **1.5×–1.8× faster than expert-written implementations**, validated end-to-end on-chip [Source 2].

### 6. Roadmap and Supply Strategy

- **Gen 1 (Jalapeño):** OpenAI reportedly plans to begin deploying it in its own infrastructure by the end of 2026 [Source 4].
- **Gen 2:** in development, heading toward tape-out, targeting better performance per watt [Source 1][Source 2].
- **Gen 3:** already "operational" (or "planned," per slide) according to Richard Ho at Hot Chips, focused on economical low-latency serving [Source 1][Source 2].

Crucially, OpenAI has stated explicitly that it intends to **continue deploying Nvidia and other partners' accelerators** alongside Jalapeño for both training and inference; the goal is supply diversity and inference economics, not a unilateral Nvidia exit [Source 3].

### 7. What This Means for AI Engineering

For builders and operators, the practical implications of the Jalapeño disclosure are:

- **Pareto-frontier thinking wins.** OpenAI's benchmark methodology (latency-vs-throughput at normalized TDP) is becoming a more useful comparison axis than peak FLOPS or theoretical HBM bandwidth [Source 1][Source 2].
- **Sliced, NUMA-style memory is the new default.** Pairing compute slices with local HBM slices, plus a dedicated low-latency collective fabric, is likely to be replicated across competitors in 2027.
- **AI-designed silicon is no longer a demo.** More than half the core authored in XLS and tuned by OpenAI's own models is now first-gen silicon already running Codex in OpenAI's labs, with kernel-level autotuners in the same boat [Source 1][Source 2].
- **Agentic inference shapes silicon.** Expect next-gen accelerators to budget silicon for small speculative models and verification, not just prefill and decode.

## Conclusion

OpenAI's Jalapeño is a credible, full-stack custom inference ASIC from a frontier AI lab. It is not a GPU, and it is not a Nvidia killer; it is a deliberately balanced, sliced, AI-designed inference platform tuned for the three-phase reality of agentic requests, scaled across 2,048-ASIC pods. The deeper story is the feedback loop: OpenAI's models designed the chip, OpenAI's models now program the kernels on the chip, and the resulting economics will make the next generation of OpenAI's models cheaper and faster to serve. For AI engineering teams, the takeaway is to start benchmarking and architecting inference along Pareto frontiers (latency vs. tokens per joule) rather than peak FLOPS, and to plan for hardware that treats prefill, draft, and verify as first-class citizens.

## Sources

1. Tom's Hardware, "Hot Chips 2026: OpenAI's Jalapeño AI ASIC unpacked" (Aug 27, 2026) — [https://www.tomshardware.com/tech-industry/artificial-intelligence/hot-chips-2026-openais-jalapeno-ai-asic-unpacked-accelerator-developed-using-ai-achieves-efficiency-and-throughput-gains-against-power-hungry-blackwell](https://www.tomshardware.com/tech-industry/artificial-intelligence/hot-chips-2026-openais-jalapeno-ai-asic-unpacked-accelerator-developed-using-ai-achieves-efficiency-and-throughput-gains-against-power-hungry-blackwell)
2. ServeTheHome, "OpenAI Jalapeno Custom AI ASIC at Hot Chips 2026" (Aug 26, 2026) — [https://www.servethehome.com/openai-jalapeno-asic-at-hot-chips-2026/](https://www.servethehome.com/openai-jalapeno-asic-at-hot-chips-2026/)
3. Quantilus, "The Custom AI Chip That Could Reshape the Inference Race" (Aug 26, 2026) — [https://quantilus.com/article/the-custom-ai-chip-that-could-reshape-the-inference-race/](https://quantilus.com/article/the-custom-ai-chip-that-could-reshape-the-inference-race/)
4. NDTV Profit, "'We Made A Chip And It Is Fast,' Says Sam Altman As OpenAI Unveils Custom Inference Chip 'Jalapeno'" (Aug 26, 2026) — [https://www.ndtvprofit.com/technology/we-made-a-chip-and-it-is-fast-says-sam-altman-as-openai-unveils-custom-inference-chip-jalapeno-11959287](https://www.ndtvprofit.com/technology/we-made-a-chip-and-it-is-fast-says-sam-altman-as-openai-unveils-custom-inference-chip-jalapeno-11959287)
