OpenAI's Jalapeno Chip Challenges Nvidia's Dominance

OpenAI's VP of Hardware, Richard Ho, discusses the company's new custom inference chip, Jalapeno, claiming it outperforms Nvidia in key benchmarks and highlighting its novel architecture.

7 min read
Richard Ho, OpenAI's Vice President of Hardware, holding up the Jalapeno chip.
Bloomberg Technology
Visual TL;DR
OpenAI's Jalapeno ChipCore
From the article 7 mentionsOpenAI is making a significant move into the hardware arena with the unveiling of its first custom inference chip, codenamed Jalapeno.
Lower Compute CostsOutcome
serving many customers with tokens much cheaper than current solutions
From the article 2 mentionsThis development marks a crucial step for the AI giant as it seeks to optimize its compute infrastructure and potentially lower costs for its rapidly growing user base.
Outperforms NvidiaEffect
From the article 4 mentionsIn a recent interview, Richard Ho, OpenAI's Vice President of Hardware, revealed that preliminary tests indicate the new chip is not only faster but also more efficient than Nvidia's GB300 system.
Challenges Nvidia DominanceOutcome
significant move into hardware arena, disrupting established market leader
OpenAI's Jalapeno ChipCore
From the article 7 mentionsOpenAI is making a significant move into the hardware arena with the unveiling of its first custom inference chip, codenamed Jalapeno.
Outperforms NvidiaEffect
From the article 4 mentionsIn a recent interview, Richard Ho, OpenAI's Vice President of Hardware, revealed that preliminary tests indicate the new chip is not only faster but also more efficient than Nvidia's GB300 system.
Novel ArchitectureContext
designed for high throughput and low latency simultaneously, an industry first
From the article 3 mentionsHe highlighted that this integrated approach, where memory and cores are tightly coupled, is a novel aspect of Jalapeno's design.
Full Stack ControlEffect
OpenAI gains control over hardware and software for better optimization
From the articleOpenAI's venture into custom silicon stems from a desire for complete control over its AI stack, from the models themselves down to the hardware.
Optimized AI InferenceContext
focus on serving AI models, not training them, for specific workloads
From the article 7 mentionsWhen questioned about not benchmarking against Nvidia's latest production system, the Blackwell, Ho clarified that the GB300 data was used because it was publicly available and optimized.
Faster Response TimesOutcome
low latency provides very fast response times for user interactions
From the article"Low latency gives you very, very fast response times.
Lower Compute CostsOutcome
serving many customers with tokens much cheaper than current solutions
From the article 2 mentionsThis development marks a crucial step for the AI giant as it seeks to optimize its compute infrastructure and potentially lower costs for its rapidly growing user base.
Challenges Nvidia DominanceOutcome
significant move into hardware arena, disrupting established market leader
Contents(5)

OpenAI is making a significant move into the hardware arena with the unveiling of its first custom inference chip, codenamed Jalapeno. In a recent interview, Richard Ho, OpenAI's Vice President of Hardware, revealed that preliminary tests indicate the new chip is not only faster but also more efficient than Nvidia's GB300 system. This development marks a crucial step for the AI giant as it seeks to optimize its compute infrastructure and potentially lower costs for its rapidly growing user base.

Jalapeno's Performance Edge

Ho explained that the performance benchmarks were conducted using InferenceX, an open-source benchmarking tool chosen for its neutrality. The results suggest that Jalapeno can achieve both high throughput and low latency simultaneously, a feat Ho described as an industry first. "High throughput means that you can serve a lot of customers with a lot of tokens much cheaper than you could if you otherwise," Ho stated. "Low latency gives you very, very fast response times. And so, you know, the agent will come back faster. You can do, you know, more coding for us, and that really, I think, is a real benefit for our products."

The full discussion can be found on Bloomberg Technology's YouTube channel.

OpenAI Says New Jalapeno Chips Outperformed Nvidia in Testing - Bloomberg Technology
OpenAI Says New Jalapeno Chips Outperformed Nvidia in Testing, from Bloomberg Technology

When questioned about not benchmarking against Nvidia's latest production system, the Blackwell, Ho clarified that the GB300 data was used because it was publicly available and optimized. He added that while Blackwell is Nvidia's latest, they aimed to use comparable published figures for their initial performance claims.

A Novel Architecture for AI Inference

The Jalapeno chip is built on an HBM4-based design, positioning OpenAI as one of the early adopters of this high-bandwidth memory technology. Ho elaborated on the chip's design philosophy, emphasizing a "blank slate approach." His team collaborated with OpenAI's research division to identify and address bottlenecks in large language models. The resulting architecture, which differs from traditional GPUs and TPUs, focuses heavily on reducing data movement and optimizing algorithms to achieve the desired low latency and high throughput. "This architecture is different from GPUs. It's different from TPUs and other chips in that nature. It really reduces the data movement. It optimizes the algorithm, and then it basically is able to get this very low latency, which matters for the agentic workloads, for example," Ho explained.

Full Stack Control and Programmability

OpenAI's venture into custom silicon stems from a desire for complete control over its AI stack, from the models themselves down to the hardware. "By having that control, we can, like, make trade-offs. We can make trade-offs and optimize it really well," Ho said. He highlighted that this integrated approach, where memory and cores are tightly coupled, is a novel aspect of Jalapeno's design. Despite concerns about the programmability of new architectures, Ho expressed confidence, noting that OpenAI was able to bring three different models up and running on the new hardware within months of receiving it, indicating a robust programming model and effective firmware optimization.

Focus on Inference, Not Training

When asked why OpenAI focused on an inference chip rather than one for training, Ho pointed to the company's immense compute needs and the growth trajectory of inference. "OpenAI has this enormous need of compute. And where that growth of compute is is really on the inference side," he stated. He acknowledged OpenAI's strong partnership with Nvidia for training hardware, allowing them to concentrate their custom silicon efforts on the inference segment, which is crucial for serving their rapidly expanding user base. Jalapeno is intended to be part of a broader fleet of inference solutions, complementing existing partnerships.

Cost Efficiency and Future Roadmap

Ho projected that Jalapeno's performance advantages would translate into lower inference costs for customers. He cited performance per watt figures between 1.8x and 4x better than existing chips, which directly impacts token costs. "We see that ultimately when this rolls out into production, customer token prices should come down," he stated. Looking ahead, OpenAI is already deep into development of its second-generation chip and has begun concept work on a third. This indicates a long-term commitment to custom silicon as a strategy to reduce overall infrastructure expenses.

The scaling of Jalapeno production involves partnerships with companies like Broadcom for chip design and Celestica for system integration. Ho acknowledged that challenges remain in areas like high-bandwidth memory and fab capacity, emphasizing the reliance on and close collaboration with partners.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.