The next AI bottleneck may be data, not model size
A former OpenAI researcher is making a forceful argument that the next major AI arms race will center on training data rather than scaling alone. According to reporting from The Decoder, Andrew Ho has left OpenAI after eight months to build a company focused on producing high-quality, specialized datasets, driven by the view that large language models still generalize poorly in economically important tasks.
The claim cuts against one of the most persistent assumptions in modern AI: that larger models, trained with more compute, will continue to broaden their competence across domains. Ho’s position is that this is not enough. In his view, many valuable forms of work are too contextual, too poorly represented, or too hard to score within existing datasets for today’s scaling playbook to reliably capture them.
That argument matters because it shifts attention from the visible competition among frontier model labs toward the less glamorous but increasingly strategic business of data creation. If Ho is right, the next durable advantage in AI may come from whoever can gather, structure, and validate task-specific examples at industrial scale.
Why current datasets may be insufficient
The source report says Ho believes the core problem is that most economically meaningful skills are barely represented in existing training corpora. Internet-scale text has been enough to create capable chatbots, code assistants, and research tools, but it may not be enough to teach models how to perform specialized work with consistent reliability.
Ho’s critique is especially pointed because it comes from inside the frontier-model ecosystem rather than from an external skeptic. He argues that even when humans can observe a “golden path” through a task, it is difficult to know whether alternate paths are also good, bad, or incomplete. That ambiguity makes many real-world skills hard to encode into neat benchmark environments.
In other words, the challenge is not simply collecting more text. It is building datasets that preserve context, outcomes, edge cases, and quality signals in domains where mistakes are expensive and where correct performance depends on tacit knowledge. That is a much harder data problem than scraping the public web.
From broad AI ambition to narrower domain execution
Ho’s thesis is reinforced in the report by work from Cambridge researcher Adam Hunt and support from Google DeepMind researchers, who argue that current systems are becoming more specialized rather than more universally capable. The implication is that some headline gains may mask a more uneven underlying picture: models can sharpen in select areas while plateauing or dulling in others.
That distinction is important for businesses trying to deploy AI into laboratories, healthcare workflows, scientific analysis, or other high-value settings. A model that performs impressively in demos but inconsistently in real environments is not enough. Reliability, repeatability, and domain-specific competence matter more than broad but shallow fluency.
If model developers are reaching a ceiling in creative problem-solving or generalized transfer, then targeted data becomes a way to push capability forward where it actually counts. That is effectively Ho’s wager: not that scaling is useless, but that its returns diminish unless paired with data designed around concrete tasks.

The first targets: bioinformatics and routine lab work
The company Ho is launching is reportedly starting with two areas. One is bioinformatics, where the article says even current models such as GPT-5.6 Sol achieve only about a 30 percent success rate on complex scientific analyses related to Ho’s earlier work at OpenAI. The second area is routine lab work, including scenarios where researchers submit images of experiments to AI systems for evaluation.
Those starting points are revealing. Both are domains where usable AI requires more than language mimicry. Bioinformatics often involves multi-step reasoning, domain conventions, scientific context, and error sensitivity. Lab workflows can depend on visual judgment, procedural understanding, and practical interpretation rather than textbook answers. These are precisely the kinds of environments where low-quality or generic data can leave models looking capable in aggregate metrics while failing in consequential use.
The report says chemistry, materials science, healthcare, and broader knowledge work are planned next. That roadmap suggests a commercial strategy built around industries where mistakes carry cost and where bespoke data could become defensible infrastructure.
A challenge to frontier-lab economics
Ho also appears skeptical of the towering valuations attached to frontier AI labs. As summarized by The Decoder, he argues that companies such as OpenAI and Anthropic are chronically unprofitable because they must keep spending heavily on new models simply to stay ahead of lower-cost rivals such as Qwen or Kimi. Whether or not one accepts that full diagnosis, it points to a growing tension in the market.
If the marginal cost of frontier scaling remains high while open or cheaper competitors narrow the gap, then proprietary advantage may need to come from something harder to replicate than model size alone. Curated datasets, private workflows, verified outcomes, and domain-specific feedback loops all fit that requirement better than raw scale. They are slower to build, harder to copy, and often more tightly tied to customer use cases.
That does not mean foundational models stop mattering. It means their differentiation may increasingly depend on the quality of the data wrapped around them. In that scenario, data companies do not sit downstream from model labs; they become a critical part of the capability stack.
What a $100 billion data market would signal
Ho’s estimate that AI labs will spend more than $100 billion on targeted data collection in the coming years is a bold forecast, but the strategic logic is easy to follow. If general web data is exhausted as a source of easy gains, then the industry needs new raw material. That raw material may be annotated scientific tasks, validated enterprise workflows, multimodal lab records, or other structured experience that current models do not see enough of during training.
Such spending would also change who benefits from the AI boom. Instead of concentrating value primarily in chipmakers, hyperscalers, and frontier-model developers, a data-centered cycle would elevate domain experts, labeling systems, synthetic-data tooling, enterprise data partnerships, and organizations capable of measuring quality in specialized environments.
The broader takeaway is that AI competition may be entering a less visible but more foundational phase. The next leap forward may not come from simply making models bigger. It may come from teaching them better, with data built for the messy, contextual, and high-stakes work that still resists automation. Ho is betting that this shift is not a side story. He is betting it becomes one of the defining markets of the next AI cycle.
This article is based on reporting by The Decoder. Read the original article.
Originally published on the-decoder.com








