OpenAI’s first custom inference chip emerges as a serious challenger
OpenAI has presented its first in-house inference chip, a processor called Jalapeño, with benchmark results that reportedly place it ahead of Nvidia’s Blackwell and Rubin platforms on several key measures. The showing matters not just because of the headline comparison, but because it suggests one of the biggest AI model developers is moving beyond dependence on merchant silicon and trying to shape the economics of inference itself.
According to the reported results from the Hot Chips conference, Jalapeño is designed only for inference. It is not a training chip, and it is not tuned solely for OpenAI’s own models. Instead, it is described as a general-purpose large language model inference accelerator meant to run models more efficiently and with lower latency. That positioning is important. Inference is where large AI systems meet users, and it is where the cost of serving chatbots, agents, and other generative applications compounds at scale.
The source text says OpenAI claims Jalapeño delivers 1.5x to 1.9x more AI work per watt at peak throughput across three tested models, along with 1.7x to 3.6x lower end-to-end latency than the best commercially available systems in those tests. For interactive workloads, the company says performance is 2.1x to 4.1x higher. If those numbers hold up in broader deployments, the chip would represent a notable first-generation result in a market where Nvidia has set the pace and where new challengers usually need several product cycles to become truly competitive.
Why these benchmarks matter
The reported tests used SemiAnalysis’s public InferenceX benchmark, with OpenAI supplying the figures and SemiAnalysis verifying some runs on site. The models named in the source were GPT-OSS 120B, Deepseek R1 670B, and Kimi K2.5 1T. On GPT-OSS, Jalapeño reportedly reached about 1,400 tokens per second per user. On Deepseek R1, it reportedly exceeded 700 tokens per second on a single concurrent request.
The source also says that at matched decoding speed, Jalapeño achieved between 54x and 104x the token throughput per kilowatt compared with the best available accelerator, depending on the model. That is a striking claim because it goes beyond raw speed and gets to the infrastructure equation that increasingly defines AI competition: how many useful tokens a company can generate for a given power budget.
Electricity, cooling, and rack density have become constraints almost as important as the chips themselves. In that environment, a design that can improve throughput per watt changes how many user sessions can be served from a fixed data center footprint. Lower latency matters just as much. For many AI products, user experience deteriorates quickly if responses slow down, especially for interactive reasoning or tool-using agents that may require many back-and-forth operations.
OpenAI is targeting inference economics, not just prestige
The deeper significance of Jalapeño is strategic. OpenAI is not merely claiming it can build a competitive chip. It is signaling that inference is valuable enough to justify vertical integration. Training still captures most of the attention in AI infrastructure, but inference is where mature products live or die financially. Every marginal gain in tokens per watt or latency can turn into lower serving costs, higher capacity, or both.
The source text notes that Jalapeño achieved its reported numbers without using techniques such as multi-token prediction or speculative decoding, while some comparison systems did use those optimizations. That means OpenAI is positioning the chip’s current results as a base rather than a ceiling. In effect, the company is arguing that there is additional headroom still available from software and system-level tuning.

That point matters because custom chips succeed only when hardware and software are co-designed well enough to turn architectural advantages into production gains. If OpenAI can pair model-serving software, compiler work, and system orchestration with a chip designed around its operational needs, it could gain leverage that goes beyond any single benchmark chart.
The caveats are as important as the claims
There are, however, reasons to read the announcement carefully. The source itself includes several constraints on how far the results can be generalized. Nvidia and AMD have already published results using larger models such as Deepseek V4 Pro and Kimi K3 that have not yet been tested on Jalapeño. That leaves open the question of how the chip performs when model sizes, context demands, and memory pressure increase beyond the reported set.
The maturity gap is another major caveat. Rubin systems are already shipping to customers, while Jalapeño reportedly has not moved beyond engineering samples. Early samples can demonstrate architectural promise, but product readiness depends on manufacturing yield, software stability, thermal behavior, deployment reliability, and supply-chain scale. A benchmark lead on a conference stage does not automatically translate into a repeatable advantage across fleets.
The source also frames total cost of ownership per token as roughly even between Jalapeño and Rubin in at least one comparison, despite the claimed efficiency edge. That is a reminder that data center economics are multi-variable. Memory, packaging, networking, deployment complexity, and utilization rates can narrow the gap between a technically superior chip and a commercially superior platform.
A first-generation chip that changes the conversation
Even with those caveats, the reported performance is consequential. First-generation chips are rarely expected to beat the market leader’s current and near-next platforms on both latency and efficiency. If OpenAI’s showing is representative, it suggests the company has moved faster than many would have expected in building specialized silicon. The source says Jalapeño was developed with Broadcom in nine months, partly using OpenAI’s own models. That short timeline, if accurate, underscores how quickly major AI developers are trying to internalize more of the stack.
The broader industry implication is that frontier model companies may no longer be content to compete only on algorithms and products. They increasingly want direct control over the hardware path that determines cost, performance, and deployment scale. Nvidia remains the incumbent with shipping systems and a deeply entrenched ecosystem, but a credible custom alternative from a top AI lab would still alter negotiating power across the market.
For now, Jalapeño should be seen as an early but meaningful signal rather than a settled victory. The most important fact is not that OpenAI has posted a benchmark win, reported though it may be. It is that the company appears to have produced a chip credible enough to be compared seriously with the industry standard. In AI infrastructure, that alone is a development worth watching closely.
This article is based on reporting by The Decoder. Read the original article.
Originally published on the-decoder.com







