Skip to main content
Ex-OpenAI researcher, training data, AI, machine learning, scaling, technology, innovation, investment.

Editorial illustration for Ex-OpenAI Researcher Sees USD 100 Billion Bet on Training Data Beyond Scaling

Ex-OpenAI Researcher Bets $100B on Training Data

4 min read

Andrew Ho spent eight months at OpenAI before deciding the company's core bet, that scaling large language models will eventually produce systems that generalize across tasks, doesn't hold up. He's now left to start a company built on a different premise: the bottleneck isn't compute or model size, it's data. Specifically, the kind of data that captures how humans actually do skilled work, from bioinformatics to routine lab procedures, none of which shows up in the text scraped from the internet that trains today's models.

Ho's new venture is building specialized datasets meant to fill that gap, and he's not shy about the price tag. He expects AI labs will collectively need to spend upward of $100 billion on targeted data collection to get anywhere close to the reliability they're promising customers. That's a striking claim from someone who spent less than a year inside one of the labs setting the pace for the industry.

His skepticism isn't isolated. Researchers at the University of Cambridge and Google DeepMind have pointed to similar trends, noting that current systems seem to be drifting toward narrow specialization rather than the broad, flexible intelligence labs keep promising.

Scaling alone won't fix this, he argues, and expects AI labs will have to spend more than $100 billion on targeted data collection in the years ahead. Ho is also skeptical of the sky-high valuations at frontier labs like OpenAI or Anthropic, which he says are chronically unprofitable because they have to keep pouring growing sums into new models just to stay ahead of cheaper rivals like Qwen or Kimi.

Why this matters

A $100 billion wager on training data is really a bet against the idea that scale fixes everything. If a former OpenAI researcher is building datasets for bioinformatics and lab work specifically because those skills are "barely represented" in what's already out there, that's a tacit admission that the industry's biggest labs have been training on a narrower slice of human expertise than the marketing suggests. For founders, this opens a real market: the boring, unsexy work of capturing "golden path" behavior in fields nobody bothered to encode.

For researchers, it's a reminder that generalization claims deserve scrutiny until someone shows the data actually covers contextual, hard-to-grade tasks. We'd push back gently on the framing that any single company can solve this at scale. Judging whether alternate approaches are "as good" as a human's golden path is a genuinely unsolved problem, not a data-acquisition problem.

Watch whether this startup's datasets actually move model performance on real lab tasks, not just benchmark scores.

Common Questions Answered

Why did Andrew Ho leave OpenAI and what is his new company's core premise?

Andrew Ho left OpenAI after eight months because he disagreed with the company's core bet that scaling large language models will eventually produce systems that generalize across tasks. His new company is built on the premise that the real bottleneck in AI development is not compute or model size, but rather the availability of high-quality training data that captures how humans actually perform skilled work in specialized domains like bioinformatics and lab procedures.

What type of training data does Ho believe is missing from current large language models?

Ho argues that current large language models lack training data that captures specialized skilled work from domains like bioinformatics and routine lab procedures. This type of data is barely represented in text scraped from the internet, which is the primary source used to train existing models, creating a significant gap in AI systems' ability to learn domain-specific expertise.

How much does Andrew Ho estimate AI labs will need to spend on targeted data collection?

Andrew Ho expects that AI labs will have to spend more than $100 billion on targeted data collection in the years ahead. He argues that scaling alone won't solve the data quality problem, making this massive investment in specialized data collection necessary for advancing AI capabilities beyond current limitations.

What is Ho's criticism of frontier AI labs like OpenAI and Anthropic regarding their valuations?

Ho is skeptical of the sky-high valuations at frontier labs like OpenAI and Anthropic, arguing that these companies are chronically unprofitable because they must continuously pour growing sums into developing new models to stay ahead of cheaper rivals like Qwen or Kimi. This spending treadmill, according to Ho, makes their current valuations unsustainable without a fundamental shift in how AI development is approached.

Further Reading

LIVE21:39Google DeepMind Demos AI Orchestrating Boston Dynamics Spot Robot