Editorial illustration for Startup Aims to Package Gaming Data for AI Learning
Startup Monetizes Gaming Data to Train AI Models
Startup Aims to Package Gaming Data for AI Learning
A player mashing buttons to dodge an enemy in a video game is, whether they know it or not, generating a small packet of physics data: an action, followed by a consequence. A British startup called Worldmodeldata is betting that this kind of raw, unglamorous gameplay, footage of ordinary people fumbling through 3D environments, could become fuel for the next generation of artificial intelligence.
The idea runs against the current playbook. Large language models got their power from text scraped off the internet, essentially an endless library of human writing. But researchers including Fei-Fei Li and Yann LeCun have started arguing that text alone can't teach a machine how to move through physical space, grip an object, or steer a vehicle. That gap has pushed part of the AI field toward "world models," systems trained to understand cause and effect in three dimensions rather than just predict the next word in a sentence.
The problem is supply. No one has assembled a data set of physical action and consequence anywhere near the scale of what fueled the LLM boom. Video games, it turns out, might be one of the few places that data already exists in bulk.
Worldmodeldata, a British startup advised by LeCun, is aiming to solve that problem by packaging controller inputs and other data collected by video game studios—an exhaust product available in massive quantities—into training datasets for world models.
Why this matters
Worldmodeldata's pitch lands at an odd moment for the industry: LLMs are running out of internet text to chew on, and the next scaling story keeps circling back to embodied, physical reasoning. Controller inputs, thumbstick twirls, trigger presses, are a strange but plausible substitute for the sensorimotor data these models lack. That's worth watching, but so is the business model underneath it.
Game studios generate this "exhaust product" as a byproduct, not a product, which means Worldmodeldata is essentially betting it can broker a supply chain that doesn't fully exist yet, one where studios agree to hand over gameplay telemetry at scale and AI labs agree it's worth paying for. LeCun's advisory role gives the idea credibility in world-model circles, but credibility isn't the same as proof the data actually improves navigation or planning in ways text can't. For founders and researchers, the real signal here isn't the dataset itself.
It's confirmation that the industry is actively hunting for anything that looks like a data type text-only training missed.
Common Questions Answered
What type of data is Worldmodeldata collecting from video games to train AI models?
Worldmodeldata collects controller inputs, thumbstick movements, trigger presses, and other gameplay data generated as an exhaust product by video game studios. This raw gameplay footage captures ordinary people navigating 3D environments and performing actions with their consequences, which the startup believes can fuel the next generation of artificial intelligence systems.
How does Worldmodeldata's approach to AI training differ from how large language models were developed?
While large language models gained their power from text scraped from the internet, Worldmodeldata is pivoting toward embodied, physical reasoning by using gaming data. This represents a shift away from text-based training as the industry recognizes that LLMs are running out of internet text to use for scaling, making sensorimotor data from games a plausible alternative.
Why is gaming data considered valuable for training world models in AI?
Gaming data provides the sensorimotor information that current AI models lack, capturing the relationship between actions and their physical consequences in 3D environments. Controller inputs and gameplay footage offer massive quantities of this embodied reasoning data that can help AI systems understand cause-and-effect relationships in physical spaces.
What advantage does Worldmodeldata have in sourcing training data compared to other AI companies?
Game studios already generate controller inputs and gameplay data as a natural byproduct of their operations, making it an abundant exhaust product available in massive quantities. This means Worldmodeldata can access large-scale training datasets without needing to create the data from scratch, unlike companies relying on scraped internet content.
Who is advising Worldmodeldata on this gaming data approach to AI development?
Worldmodeldata is advised by LeCun, a prominent figure in artificial intelligence research. His involvement suggests credibility in the startup's approach to using gaming data for training world models.
Further Reading
- Worldmodeldata raises £7m to accelerate gaming data for AI training - UKTN
- Worldmodeldata lands £7M to turn gaming data into AI training - Tech.eu
- Worldmodeldata raises £7M seed round - The SaaS News
- Origin Lab raises $8 million to glean AI training data from game worlds - GamesBeat
- General Intuition's $2.3B bet that video games can train AI agents for the real world - TechCrunch