OpenAI and Broadcom have officially revealed Jalapeño, OpenAI's first custom-designed inference processor. This accelerator is engineered specifically for the demands of large language models (LLMs) and represents a significant step in OpenAI's ambition to build out its entire infrastructure stack.
The collaboration aims to deliver a multi-generation compute platform designed to make advanced AI faster and more accessible. Early tests indicate that the first-generation Jalapeño chip will offer substantially better performance per watt compared to existing state-of-the-art accelerators. This new OpenAI Broadcom LLM chip is built from the ground up, considering the needs of current and future LLMs across the industry.
Accelerated Development and Full-Stack Vision
Developed in an accelerated nine-month tape-out process, the chip's design was informed by OpenAI's deep understanding of LLM fundamentals and its roadmap of models and serving systems. Broadcom and Celestica were key partners in industrializing the platform, handling chip implementation, board design, rack system integration, and scalable production.
Jalapeño is designed for flexibility, intended to work with a wide range of LLMs. Engineering samples are already running workloads, including GPT‑5.3‑Codex‑Spark, at target specifications. The architecture focuses on reducing data movement and optimizing compute, memory, and networking resources to achieve high utilization rates.
"The world is moving to a compute-powered economy," stated Greg Brockman, President and Co-Founder of OpenAI. "Jalapeño is part of our long-term full-stack infrastructure strategy to make compute more abundant."