the feed MANY MINDED · THE BRIEF
KNOWLEDGE · forward · impact 2/5 · 2026-08-26 · OpenAI

OpenAI's Jalapeño chip outperforms Nvidia's Rubin in key benchmarks but remains in engineering samples

OpenAI's engineering sample chip Jalapeño outperforms Nvidia's Rubin in key benchmarks but remains in engineering samples

OpenAI demonstrated its first custom inference chip, Jalapeño, at Hot Chips, showing it outperforms Nvidia's Rubin in critical benchmarks for AI workloads. The chip delivers 1.5x to 1.9x more AI work per watt and 1.7x to 3.6x lower end-to-end latency than Rubin across major models like GPT-OSS 120B and Deepseek R1 670B. It achieves 1,400 tokens per second per user on GPT-OSS 120B and over 700 tokens per second per concurrent request on Deepseek R1 670B—while using 54x to 104x less power per token than current accelerators. Jalapeño was designed without multi-token prediction techniques, using OpenAI's own models during development, and is currently in engineering samples with no customer shipments.

The chip's efficiency gains stem from its architecture optimized for inference workloads without training AI models. By avoiding techniques like speculative decoding that some competitors use, Jalapeño achieves higher token throughput per kilowatt than Rubin despite Nvidia's optimizations. This could lower the cost of AI services globally if scaled, as OpenAI claims its design complements existing partnerships with Nvidia, AMD, AWS, and others rather than replacing them.

For affordability, Jalapeño's performance could expand access to AI services where compute costs are a barrier—particularly in regions with limited infrastructure. However, it remains in engineering samples, has not yet shipped, and benchmarks exclude larger models like Deepseek V4 Pro that Nvidia has published. Real-world testing with actual users will determine if these gains translate to cheaper, more accessible services beyond lab conditions. The next step is November 2025 fabrication, after which OpenAI will need to validate performance in production environments.

Source: The Decoder