28 days ago
TechCrunch Aug 25, 2026

OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show

OpenAI unveiled detailed performance benchmarks for its new Jalapeño chip at the Hot Chips conference on August 25, 2026. Developed in collaboration with Broadcom, the chip excels in AI inference workloads, showing significant improvements over existing processors like Nvidia’s Blackwell system. Tested on SemiAnalysis' InferenceX benchmark, Jalapeño demonstrated higher throughput per kilowatt and processed more tokens per user, making it both faster and more power-efficient in handling AI tasks. OpenAI plans a phased rollout starting with limited volumes by the end of 2026 and broader deployment in 2027.

Jalapeño was designed with a full-stack approach, integrating AI models, chips, memory, and software to optimize inference performance. OpenAI specifically targeted common bottlenecks in prefill and communication phases of inference, minimizing data movement and keeping model state, including KV cache, local. This design choice helps reduce latency while maintaining high efficiency, enabling the chip to handle large-scale AI workloads rapidly and cost-effectively.

OpenAI’s head of hardware, Richard Ho, emphasized the chip’s ability to serve numerous AI users simultaneously with low latency and superior energy efficiency compared to current state-of-the-art processors. While the chip’s initial benchmark leads Nvidia’s best offerings today, Ho acknowledged that technology may advance by the time Jalapeño reaches full deployment. Nonetheless, OpenAI is committed to making Jalapeño a multigenerational platform to continue adapting AI hardware alongside evolving models.

By building Jalapeño as part of a comprehensive hardware-software stack, OpenAI aims to better control AI inference performance and cost. This strategy reflects broader trends in AI development where owning the entire AI infrastructure—from models to chips—allows companies to optimize speed, energy use, and scalability. Jalapeño represents a major step forward in OpenAI’s infrastructure, underpinning the company’s ongoing efforts to improve AI accessibility and responsiveness at scale.

0
0 Read source
Share this post
Facebook Twitter LinkedIn

Discussion

0 comments

No comments yet

Start the discussion with a take, question, or market read.