OpenAI Debuts Jalapeño Inference Chip to Break Nvidia's AI Monopoly
Co-developed with Broadcom in just nine months, the 3-nanometer ASIC slashes power consumption while delivering a massive performance boost for agentic workloads.

Hardware development is notoriously sluggish, measured in years and massive capital expenditures. OpenAI just tossed that playbook out the window. On August 25, the AI heavyweight officially entered the silicon arena with "Jalapeño," a custom-built Application-Specific Integrated Circuit (ASIC) co-developed with Broadcom that took a mere nine months from tape-out to reality.
The End of the Throughput-Latency Compromise
Traditionally, chip architecture forces a brutal choice: you either optimize for throughput to process massive amounts of data, or you optimize for latency to deliver lightning-fast individual responses. Jalapeño's bespoke design breaks this compromise. Built on TSMC's cutting-edge 3-nanometer process, it packs ultra-fast HBM4 memory delivering 15.4TB/s of bandwidth per package.
What makes this genuinely remarkable is how little power it requires to achieve this performance. Jalapeño operates at a Thermal Design Power (TDP) of just 700 watts. For context, Nvidia's latest Rubin chips devour anywhere from 900 to 1,150 watts. According to the independent InferenceX benchmark suite, OpenAI's silicon delivers 1.5 to 1.9 times more AI work per watt than comparable commercial systems.
This efficiency is exactly what makes complex, continuous chain-of-thought reasoning possible without bankrupting developers. Hock Tan, CEO of Broadcom, noted that early lab tests indicate Jalapeño's operating costs could be approximately 50% lower than those of common AI GPUs under typical workloads.
Machines Designing Machines

The tech community's shock isn't just about the benchmarks; it's about the unprecedented timeline. A 16-month cycle from team formation to manufacturing tape-out is practically unheard of in modern silicon engineering. OpenAI achieved this impossible schedule by deploying its own frontier models to co-design the architecture.
The biggest hurdle for any new chip is the software stack. Nvidia's CUDA platform is famously entrenched, acting as a massive moat against competitors. OpenAI started from zero but overcame this barrier by having its AI translate and optimize the necessary code for Jalapeño's unique architecture.
“The world is moving to a compute-powered economy. Jalapeño is part of our long-term full-stack infrastructure strategy to make compute more abundant.”— Greg Brockman
Interestingly, this hardware isn't locked into a walled garden. While designed in-house, Jalapeño achieved top-tier performance on massive open-source models like DeepSeek R1 and Kimi K2.5, proving its viability as a highly generalized inference engine.
What people are saying
“The real breakthrough in the NVIDIA SANA team’s Sol Engine work on MiniMax H3! By splitting generation into a 4-step low-res H3 draft and a 3-step LTX refinement pass at target resolution with Sol-Attn, they’ve crushed 10s 768p latency on a single GB200 from 414s down to 14.93s”
“Xiaomi just showed its AI Cube Prototype and this could become a serious GB10 competitor from China 👀 - 3 custom chips: Xring O3, O100, D100 - 200 TOPS NPU - 1.22 TB/s AI memory bandwidth - Up to 160GB unified memory - 150W sustained power - 120B models running locally Xring”
“The first CPU built for agents is now at work at massive scale. @SpaceXAI is deploying NVIDIA Vera to power agentic AI — faster agents, GPUs fully utilized, and a single architecture from Earth to orbit. Read more ⬇️”
OpenAI Jalapeño vs Nvidia
More stories






