Prompt and Model
Models

Worldmodeldata Trains AI World Models on Video Game Data

British startup Worldmodeldata is licensing nearly one million hours of video game telemetry to train AI world models, aiming to solve the data bottleneck

British startup Worldmodeldata is licensing nearly one million hours of video game telemetry to train AI world models...

Worldmodeldata has licensed almost one million hours of video game controller and telemetry data to train AI models that understand real-world physics. The British startup, advised by celebrated researcher Yann LeCun and led by CEO Rhea Loucas, acts as a broker curating and organizing this data for AI labs.

World models are designed to understand physical cause and consequence, spatial reasoning, and object behavior. They require visual data paired with precise action metrics, but there is no equivalent to the internet's vast text corpus for physical interactions. Worldmodeldata packages controller inputs and other exhaust data generated by modern gaming engines, which log telemetry at sixty frames per second. This includes visual output, camera angles, and collision metrics. The company normalizes this disparate telemetry into a standardized format for machine learning pipelines.

Founders see games as a path to scalable world models

CEO Rhea Loucas argues video games offer the abundant, diverse experiences needed to train world models. She believes this data could become the majority of training material, potentially triggering a major 'GPT moment' for the field. "Why don’t we take the vast, abundant, diverse experiences from video games, and teach AI?" Loucas said. The company's theory is that video game data provides the necessary quantity and variety to teach AI about rare corner cases. Worldmodeldata also envisions a future where individual players could be compensated for the kinetic data they generate while playing.

Experts debate suitability of game data for physical AI

Not everyone shares this optimism. Ming-Yu Liu, who leads world model development at Nvidia, suggested game data is better suited for training models to generate hyperrealistic video or 3D environments, not physical robotics. He expressed caution regarding utility for tasks requiring fine-grained motor control. "I would be more conservative on using video game data for manipulation. The physics for manipulation is much more involved," Liu stated, noting developers often take shortcuts to create realism.

Xiatian Zhu, an associate professor specializing in AI at the University of Surrey, described video games as essentially coarse simulators. He argued their approximations remain too crude for the exacting demands of real-world robotics. This debate highlights a fragmentation in the race for artificial general intelligence. While some firms like Niantic and General Intuition harvest spatial data from their own platforms, others rely on synthetic data generation. For researchers like Stanford's Fei-Fei Li, who champions spatial intelligence, true machine intelligence requires a foundational understanding of the physical universe.

VC backing and broader context in spatial AI

The shortage of suitable training data is considered one of the largest bottlenecks to progress. Nicole Fraenkel, a partner at VC firm Khosla Ventures, emphasized the critical importance of corner cases for physical systems. The corner cases are the ones to actually get right, Fraenkel said, noting the high cost of error with vehicles, drones, or robots. She pointed out that repetition alone won't capture the world's disorder for machine operation. While the hypothesis that performance improves with dataset size-similar to large language models-is yet to be fully tested, Khosla Ventures has invested in similar spatial AI initiatives. Until world models reach their breakthrough moment, many ideas remain on the table. Worldmodeldata aims to create avenues for individual players to be compensated for the kinetic data they generate while playing.

Topics

#Models

Related coverage

More from Models