Gameplay video paired with the ground truth a world model actually needs — per-frame camera pose & intrinsics, metric depth, 3D object tracks, and full player action, all aligned on one timeline. Produced on a proprietary industrial engine (UE5-class).
Metric depth (meters, float32) and oriented 3D boxes come straight from the engine, with exact intrinsics and 6-DoF pose per frame. No monocular-depth or structure-from-motion error to inherit — the world state itself, not a guess from pixels.
World state can be snapshotted and re-run: push the bottle off the table, remove an object, change an action. The same scene evolves different ways — the exact signal a world model needs to learn action results, causality, and counterfactuals.
Map a model's predicted action back into the engine, regenerate the matching ground truth, and score its output on physical plausibility and temporal consistency. The pipeline that produces training data also produces an eval harness. 95%+ QC; billed on qualified data only.
Gameplay video that carries its own 4D ground truth — per-frame camera pose, metric depth, and 3D object tracks, with the player’s every input, all measured by the engine on one clock rather than estimated from pixels.