OpenAI's New Jalapeño Accelerator Outperforms Nvidia in Recent Benchmark Tests

OpenAI’s Jalapeño achieved 1.5x–1.9x higher AI work per watt and 1.7x–3.6x lower end-to-end latency than Nvidia’s GB300 across tests using DeepSeek R1, Kimi K2.5 1T, and GPT-OSS-120B.
Physically, Jalapeño is built as a 128-chip configuration capable of about 1.7 exaFLOPS with 27 TB of HBM memory, underscoring the emphasis on memory bandwidth for inference performance.
OpenAI describes Jalapeño as an inference-focused ASIC developed with Broadcom, marketed as delivering the 'best of both worlds' with lower latency and higher throughput, while explicitly not serving training workloads.
OpenAI plans to roll out Jalapeño in small volumes by year-end with a ramp in 2027, and while it will expand capacity, it does not intend to replace its entire compute stack and will continue to rely on existing partners like Nvidia and AMD; a second-generation design and a third version are already in development.
OpenAI indicated it has no plans to sell Jalapeño chips to third parties for now, citing a substantial internal demand for compute and the scale of its needs.
OpenAI unveiled its first custom AI chip, Jalapeño, claiming it outperforms Nvidia's GB300 in two critical areas: power efficiency and response speed. Yahoo News reports the inference-focused accelerator delivers 1.5x to 1.9x more AI work per watt and 1.7x to 3.6x lower end-to-end latency across multiple large language models. OpenAI will begin small deployments by year's end, with a full ramp expected in 2027.
The company built Jalapeño using 128 chips per configuration, generating about 1.7 exaFLOPS of computing power with 27 terabytes of memory, Data Center Dynamics reports. OpenAI developed the chip with Broadcom and emphasized it serves only inference workloads—the task of answering user questions—not training new AI models. The company plans to continue working with Nvidia, AMD, and Broadcom rather than replace its entire hardware stack.
Building custom silicon lets OpenAI cut costs and reduce reliance on Nvidia. The company faces huge demand for AI inference—each ChatGPT response requires fast, efficient processing. Head Topics notes OpenAI presented Jalapeño at Stanford's Hot Chips semiconductor conference, signaling the broader industry shift toward in-house accelerators. Creating dedicated inference chips lets OpenAI optimize for speed and power, addressing its own bottlenecks.
Jalapeño tested against Nvidia's GB300 using three demanding models: DeepSeek R1, Kimi K2.5 1T, and GPT-OSS-120B. Benzinga reports OpenAI's chip delivered substantially better efficiency and speed. The 1.5x to 1.9x jump in AI work per watt means Jalapeño accomplishes more tasks using less electricity. The 1.7x to 3.6x latency drop means responses arrive faster—critical for live user interactions where milliseconds matter.
Jalapeño handles inference—running existing AI models to generate answers. It does not train new models, a compute-intensive task requiring different hardware. Yahoo News confirms OpenAI will keep using external GPUs and partners for training work. Jalapeño's narrow focus allows OpenAI to optimize every aspect for inference speed and efficiency. The company has no plans to sell Jalapeño to other organizations yet, citing its own massive internal compute demand.
OpenAI is already designing Jalapeño's successor and has begun early planning for a third-generation chip. Data Center Dynamics reports Jalapeño will start with small volumes by late 2024, ramping through 2027. The staggered timeline reflects the complexity of custom silicon—designing, manufacturing, and deploying new chips takes years. By planning multiple generations now, OpenAI positions itself to steadily reduce inference costs and speed.
Publishers
17
Articles
24
Reach
41