OpenAI Puts Most of Its Research Into Models That Don't Exist Yet
OpenAI is devoting the overwhelming majority of its research capacity to models that have not been built. Boris Power, the company's Head of Applied Research, puts the figure at 80 to 90 percent of all research effort going toward GPT-7, GPT-8 and the generations beyond them. The reasoning is straightforward: that is where the company believes "most of the value" will ultimately be created.
The disclosure, made during a session at the Fellows Forum, reframes how outsiders should read OpenAI's steady stream of mid-cycle model releases. Those updates are real products that customers use every day, but they are not where the company's research attention is concentrated.
Power's framing suggests an organization that treats each numbered generation as a platform shift rather than a product iteration — a step change that resets assumptions about what the technology can do and where the returns will come from.
The scale of that commitment is unusual even by the standards of frontier AI labs. Allocating the clear majority of research resources to systems that are several generations away from shipping implies a belief that the biggest economic and capability gains are not available through refinement of what already exists.
Incremental Updates Are a Deliberate Short-Term Bet
Between major generations, OpenAI ships smaller improvements. Power pointed to the jump from GPT-5.1 to GPT-5.2 as a representative example. Work of that kind leans heavily on specialized training data — targeted datasets that sharpen a model's behavior in specific areas without altering the underlying approach.
Internally, Power said, these incremental releases are regarded as "extremely shortsighted." That is not so much a criticism as an admission of strategy: the company knows they are not the long game, and it ships them anyway.
The justification is speed. Frequent updates let OpenAI iterate, gather signal and learn faster in the present, even though the approach will not carry the company where it wants to go over the long term. In other words, the short-term bets help fund the learning that informs the long-term ones.
- Most research capacity: aimed at GPT-7, GPT-8 and beyond.
- Mid-cycle updates: driven by specialized training data.
- Internal view: these updates are useful, but strategically narrow.
- Long-term payoff: expected to arrive with new model generations.
The Real Leap Comes With a New Generation
Power described the generational jump as the moment when "everything else just works a lot better." Capabilities that were brittle or partial in one generation tend to cohere in the next, producing gains that no amount of fine-tuning on the previous model can replicate.
The distinction is between sharpening a model on narrow tasks and raising the ceiling on everything at once. Specialized training data can improve how a model handles particular categories of requests; a new generation can improve how it reasons, follows instructions and copes with ambiguity across the board.
That dynamic carries an unusual operational cost. After every step up, OpenAI has to relearn where to invest for quick wins. Techniques that once produced meaningful gains can become redundant when the base model improves on its own — a moving target that forces the company to rebuild its playbook with each generation.
Onboarding, Not Model Quality, Is the Bottleneck
Perhaps the most consequential part of Power's account concerns what he sees as the primary obstacle facing AI assistants today. It is not the quality of the models. It is onboarding.
Most ChatGPT users, he said, do not know what they can actually do with AI. The gap is not one of capability but of discovery: the models can already handle a wide range of tasks that many users never think to attempt. Closing that gap is a design and product problem as much as a research one.
This is a notable claim from someone whose job sits close to the research frontier. It suggests that the limiting factor on real-world value is less about how smart the next model is and more about whether people can figure out what to ask it to do. A more capable system that users underutilize delivers less than its potential.
Power's prescription for future models is that they should become better at surfacing what is possible and at anticipating what a user needs before being asked. That points toward assistants that propose actions, suggest workflows and reduce the burden of figuring out where to start.
How Each Generation Changes the User Experience
Power sketched the evolution of usability across recent OpenAI models in concrete terms:
- GPT-4 demanded careful prompting — users had to learn the craft of asking.
- GPT-5 became easier to work with, but still required a great deal of feedback to stay on track.
- GPT-6 behaves more like a capable colleague: you can hand it a goal and let it work.
The trajectory he describes is a shift from instructing a tool to delegating to a collaborator. Each step reduces the amount of user expertise required, which in turn should broaden who can extract value from the system.
What This Says About OpenAI's Roadmap Logic
Taken together, Power's comments outline a two-track strategy. One track keeps the current product competitive through frequent, data-driven updates. The other track — where most of the research budget flows — is aimed at generational resets that make the previous track's optimizations obsolete.
That division also explains why the company treats each new generation as the decisive event. If the biggest gains come from the reset rather than the refinement, then research priorities should follow the reset, even when the intermediate releases are what customers actually use today.
The unresolved question is the one Power himself raised: whether the value already built into today's models can reach the people who are not yet using them well. By his own account, that is less a matter of making the models smarter and more a matter of making them legible to everyone else.
For now, the company appears comfortable with the trade-off. It will keep shipping the small updates while staking the bulk of its research on the generations that have not arrived.
This article is based on reporting by The Decoder. Read the original article.
Originally published on the-decoder.com







