OpenAI’s ‘Jalapeno’ Chip Debut Shows a Competitive First-Generation Design

OpenAI unveiled its custom inference chip, “Jalapeno,” at Hot Chips 2026 on August 25, 2026. Semiconductor research firm SemiAnalysis was invited to OpenAI’s lab and validated the results using its InferenceX benchmark suite. Key facts include:
- Release timeline: Publicly unveiled on August 25, 2026
- Chip status: Engineering samples are ready; production is expected to ramp gradually in 2027, with most output concentrated toward late 2027
- Power consumption: 700W TDP
- Performance ranking: Across multiple open-source models, it beat every NVIDIA, AMD and Google chip SemiAnalysis was able to test
- Partner: Co-designed with Broadcom
- Deployment plan: OpenAI plans to work with neocloud vendors, first gathering reliability data before scaling up
Benchmark Results: Efficiency and Latency Leadership

According to SemiAnalysis’ on-site validation, Jalapeno showed strong results on three open models: GPT-OSS 120B, DeepSeek R1 670B and Moonshot AI’s Kimi K2.5 with 1 trillion parameters.
- Per-watt throughput: 1.5-1.9x that of NVIDIA GB200 NVL72 and GB300 NVL72 rack systems
- Inference latency: End-to-end latency was 1.7-3.6x lower than NVIDIA’s recorded best results for GB200 NVL72 and GB300 NVL72; in ultra-low-latency scenarios, it was 2.1-4.1x faster
- Peak throughput comparison: Under GB300’s fastest-output setting, per-kilowatt throughput was up to 8.6-104.3x higher
- Per-user throughput: Under low concurrency, GPT-OSS and Kimi K2.5 reached about 1,400 tokens/second/user, while DeepSeek R1 exceeded 700 tokens/second/user at concurrency 1
- Correctness: GSM8k results matched NVIDIA chips
One caveat matters: these results were based on single-token prediction (STP), without speculative decoding or prefill/decode separation. The Blackwell comparison data used multi-token prediction (MTP). SemiAnalysis said that if Jalapeno is compared with GB300 running MTP, its peak energy-efficiency lead narrows to about 1.5x.
The benchmark numbers were provided by OpenAI, and SemiAnalysis validated the InferenceX benchmark process on-site. However, the team did not run the full suite and did not test AgentX, its preferred benchmark for long-context, multi-turn dialogue workloads. The data therefore points to strong engineering-sample performance under specific conditions, not a final production verdict.
Technical Breakthrough: AI-Augmented Design and Rapid Iteration
Jalapeno design began in mid-2024. The chip completed tape-out in November 2025, including CoWoS package design, and powered on about 3 months later. OpenAI produced A0 stepping results within 9 months. The second B0 stepping has entered tape-out/fab and is expected to improve performance per watt by about 25%.
Key specifications:
- B0 stepping uses a single compute die on TSMC N3P, paired with an I/O chiplet on N3E
- MXFP4 compute reaches 13.4 PFLOPS
- Six HBM4 stacks deliver 216GB of capacity and 15.4TB/s of bandwidth in one package; the source describes this as the highest among shipped or near-shipping accelerators, with Samsung likely the supplier
- The software stack uses Gluon, a Triton-based language that keeps the SPMD programming model while exposing lower-level abstractions
AI-assisted development: With help from Codex and GPT-Astra, the SIMD unit area was reduced by 8% and the matrix-engine area by 10%. Some AI-generated kernels were 1.5-1.8x faster than kernels written by human experts.
OpenAI also used Codex to port Doom to the chip, where it ran at 36 FPS.
System Architecture: Vindaloo Rack and Scaling Plan

The rack system built around Jalapeno is called “Vindaloo.” Its configuration is:
| Metric | Value |
|---|---|
| Chips per rack | 128 Jalapeno chips |
| Host architecture | Katsu CPU host rack + Chana switch rack |
| Two-rack power | About 160kW |
| Expansion capability | Up to 16 racks and 2,048 chips per scale-up domain |
| Next target | 100MW scale |
Important note: OpenAI does not currently plan to replace all of its existing chip deployments with Jalapeno. It will continue working with NVIDIA compute partners while also developing second- and third-generation chips.
Industry Impact and Deployment Guidance

NVIDIA’s stock rose 2.19% despite the news. The source points to two possible reasons: NVIDIA launched the Jetson Orin Nano 2 robot computer on the same day, and OpenAI said it would not abandon NVIDIA, would still rely heavily on NVIDIA chips for training, and has financing-backed cooperation with NVIDIA.
Reader recommendations:
- Inference-heavy developers: If ultra-low latency and energy efficiency are priorities, watch OpenAI’s ecosystem access and availability after the 2027 production ramp
- Enterprise IT decision-makers: It is still too early to make procurement assumptions; Jalapeno remains an engineering sample, and production stability plus real-world workload performance still need validation
Final Thoughts
Jalapeno’s fast progress signals a new phase in cloud infrastructure competition. The source also notes an irony: GPT 5.6 Sol, running on NVIDIA GPUs, helped design a chip that may challenge CUDA’s moat. AI itself is becoming an accelerator for chip-design iteration.
NVIDIA’s positive stock reaction also suggests that the market is not pricing the story purely as a replacement threat. As long as OpenAI still relies heavily on NVIDIA for training compute, ecosystem cooperation and near-term supply dynamics remain just as important.
