Game studios have been sitting on a goldmine of player actions. Now, a British startup wants to turn those digital moves into the next big leap for artificial intelligence. Worldmodeldata is collecting the raw, messy gameplay of millions-every missed jump, every awkward turn. Their goal is simple. Feed this chaos to AI world models and help machines finally understand how things move and collide in the real world. So far, even the best language models have struggled with that.
GameHorizon recently released a dataset with 5,000 hours of AAA gameplay from 21 titles, collected by 100 expert players, including temporally aligned videos and player actions.
Worldmodeldata’s answer is to license controller inputs and gameplay logs straight from studios. There’s a mountain of this data. CEO Rhea Loucas says the company has locked in nearly 1 million hours already. She won’t say which studios. The company acts as a broker, sorting and packaging the data so AI labs don’t have to chase down every developer. Others are moving in too. General Intuition and Niantic are building their own gameplay data pipelines, collecting from their own games and platforms.
Loucas is blunt: “There are millions of great games, and they are more and more similar to the real world. Why don’t we take the vast, abundant, diverse experiences from video games, and teach AI?” The idea is that the scale and variety of game data can help world models learn not just the basics, but also the rare, unpredictable moments-corner cases-that trip up machines in the wild. Nicole Fraenkel, partner at Khosla Ventures (which backs General Intuition), says, “The corner cases are the ones to actually get right. The cost of error with a car, plane, drone, factory forklift, or autonomous quadruped is very high.”
GameHorizon evaluated 47 different models using over one million model invocations, demonstrating that gameplay and embodied-interaction data are now being used at substantial scale for benchmarking agent capability.
Academics are wary too. Zhu points out that video games are just simulators with some physical grounding. Their worlds are “very coarse, approximate.” The risk is clear. Models trained on these inputs may stumble when they meet the messy, tactile demands of the real world.
Still, the need for scalable training data keeps pushing the field. Worldmodeldata is already eyeing the next step: licensing from individual players. Gamers could get paid for their digital skills. Loucas thinks video game data will become the backbone of world model training. Real-world data will come later, for fine-tuning. She sees a “GPT moment” ahead, when world models finally become useful for real-world tasks.
The stakes are high. Whoever cracks scalable world model training could unlock new levels of automation, from robotics to hyperrealistic content. Fraenkel sums it up: “There are many paths to the promised land. The truth is, the jury is still out on which one is going to work best.”
While the debate over simulated versus real-world data continues, the flood of video game actions into AI training marks a sharp turn. This shift lines up with reported earlier experiments showing that oddball data sources can shake up AI playbooks. If Worldmodeldata’s bet works, the next AI breakthrough may come not from a lab, but from the muscle memory of millions of gamers. The race is on. It’s not just about algorithms anymore. It’s about who controls the most diverse, actionable data. For creators and publishers, this could mean a future where digital play and real-world automation blur. Even the smallest data point could reshape the AI landscape.
Worldmodeldata, founded in the UK, has moved fast. Its licensing of nearly 1 million hours of gameplay is one of the biggest deals so far. As demand for world model training data grows, this approach could set a new standard for how digital experiences are turned into fuel for AI.