OpenAI's Jalapeño ASIC Challenges Nvidia's Dominance with Superior Inference Performance
OpenAI's newly unveiled Jalapeño ASIC, a custom inference chip co-developed with Broadcom, claims significant performance per watt and latency advantages over Nvidia's flagship GB300 systems, as presented at Hot Chips.
✨ This content was summarized and interpreted by AI; it may contain errors — please verify accuracy with the original sources. Learn more
Listen to this story

OpenAI has unveiled its Jalapeño ASIC, a custom inference chip co-developed with Broadcom, presenting benchmarks at Hot Chips on Tuesday that claim significant performance per watt and latency advantages over Nvidia's flagship GB300 systems. The 700-watt Jalapeño processor reportedly delivers between 1.5 and 1.9 times more AI work per kilowatt and achieves 1.7 to 3.6 times lower end-to-end latency compared to Nvidia's GB200 and GB300 configurations, based on tests conducted using the SemiAnalysis InferenceX benchmark suite. These compelling figures were observed across a range of large language models, including OpenAI's own GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI's Kimi K2.5 1T, with the most substantial leads occurring at low-latency operating points, where Jalapeño demonstrated 8.6 to 104.3 times higher throughput per kilowatt for single-token prediction. While the Nvidia accelerators used for comparison were rated at 1,200W and 1,400W, Jalapeño's measured sustained power remained at or below 550W during testing.
This strategic move marks OpenAI's entry into custom silicon, positioning it alongside tech giants like Google, Amazon, Meta, and Microsoft, all of whom are investing heavily in proprietary chips to power their burgeoning AI infrastructures. The imperative behind this shift is clear: to mitigate reliance on Nvidia, which currently commands approximately 70% of the AI chip market, and to gain greater control over the hardware-software stack. For OpenAI, this vertical integration is critical for optimizing performance, reducing escalating operational costs — particularly electricity, a dominant expense in data centers — and ensuring the scalability and reliability of its AI models for both consumer products and enterprise services. The Jalapeño, designed as a "blank-slate" for modern LLM inference and a generalized chip capable of running any model with ease, is not a mere adaptation but a purpose-built accelerator.
The co-development process with Broadcom, a pivotal enabler in the custom ASIC ecosystem, was remarkably swift, moving from initial design to manufacturing tape-out in an astonishing nine months, a pace some consider the fastest ever achieved in high-performance semiconductor development. This speed was partly attributed to OpenAI's innovative use of its own AI models to accelerate portions of the chip's design and optimization. Broadcom's expertise in silicon implementation, networking, and connectivity, including its Tomahawk networking silicon, is instrumental in bringing this multi-generation platform to large-scale production. The chip itself is fabricated on TSMC's advanced 3nm process, underscoring its cutting-edge nature.
The implications for the broader AI and semiconductor industries are profound. Nvidia's Grace Blackwell Superchip (GB200) and its rack-scale GB200 NVL72 system, featuring 72 Blackwell GPUs and 36 Grace CPUs interconnected by NVLink-C2C, represent a monumental leap in computing power, with the GB200 NVL72 boasting a thermal design power (TDP) of 2700W and delivering 1.44 EFLOPS of FP4 compute per rack. While Nvidia's hardware remains unchallenged in the computationally intensive domain of AI model training, OpenAI's Jalapeño directly targets the inference workload, which is experiencing an "inference inversion" in 2026, where the volume of tokens generated by deployed models eclipses the compute requirements for training. This shift in demand dynamics creates a fertile ground for specialized inference ASICs.
The market for custom AI ASICs is experiencing explosive growth, projected to expand from USD 43.8 billion in 2026 to an estimated USD 308.3 billion by 2035, driven by the insatiable demand for higher performance, lower power consumption, and reduced inference costs. This burgeoning market, with Broadcom alone targeting $100 billion in annual AI chip revenue by 2027, signifies a structural change in how AI compute is procured and deployed.
Looking ahead, OpenAI plans to begin deploying Jalapeño in its data centers in small volumes by the end of 2026, with a significant ramp-up anticipated through 2027. This initial rollout is merely the first step in a "multi-generation compute platform," with a second-generation Jalapeño already slated for tape-out in the coming months and a third generation in conceptual stages. While OpenAI acknowledges Nvidia as a "really good partner" and will continue to utilize Nvidia GPUs for certain workloads, particularly training, the introduction of Jalapeño undoubtedly intensifies competitive pressure on Nvidia, especially concerning pricing power in the inference segment. The true test will lie in sustained, independent verification of these benchmarks and the seamless integration of Jalapeño into OpenAI's production environments, but the initial announcement signals a pivotal moment where a leading AI developer is actively shaping the silicon landscape to meet its own advanced compute needs.