Ego-centric world models

What 30,000 hours of ego-centric video does not teach

Agent and object fidelity

Nano baseline · select a scale

Trained on 30k hours

OverlapGT onlyPredicted only
GT
Nano · 30k h

Ego-centric data scaling reaches high agent fidelity, but the effect of the agent's actions on the world lags far behind.

The hands land almost exactly where they belong — the paper they are folding does not.

01 · Dataset

30,012 hours of first-person video.

We train all models on a large-scale ego-centric dataset of everyday human manipulation: 1.15M first-person clips totalling 30k hours (3.24 × 109 frames at 30 fps, mean clip length 94 s), each recording a manipulation activity from a head-mounted camera.

Mecka

We thank Mecka for their support with the dataset.

01 / Duration
First-person video
02 / Recordings
Clips · mean 94.3 seconds
03 / Diversity
Unique tasks
04 / People
Contributors
02 · Findings

More data gets the hands right, but not the objects they handle.

Our models see one first-person frame and the motion of the person's arms and hands, then generate the next 16 frames (1.6 s). We score the agent (the person's hands) and the objects they handle separately, on clips recorded with different people, places and cameras than training. To see what more data would bring, we fit a scaling curve to each result and extend it beyond 30k hours (dotted lines). We then use skeleton conditioning as an instrument for locating the object limit. It additionally supplies these poses as a skeleton projected onto the image, so the model no longer has to infer where the hands appear in the frame, and the hands reach their limit with far less data.

Metric · SCS (structural consistency)

SCS segments and tracks key scene elements through the predicted and ground-truth videos and averages mask IoU, capturing whether the provided actions are accurately reflected in the generated world.

How SCS is computed Example
Element
Trained on
Ground truthGT mask
Predictionpredicted mask
Mask overlayOverlapGT onlyPredicted only
Frame 1 / 16–IoU = overlap ÷ (overlap + GT only + predicted only)
Clip SCS–mean IoU over future frames 1–16 (dashed line)
1 · Segment & track
Key scene elements, the hands and each manipulated object, are segmented and tracked through the predicted and ground-truth videos. Ground-truth masks are annotated with SAM 2 and manually verified.
2 · Compare masks
In every frame, the predicted mask is compared with the ground-truth mask by IoU: their overlap divided by their union.
3 · Average
Mask IoU is averaged over future frames 1–16 and reported separately for the hands (agent) and the manipulated objects. Higher is better.

Because mask tracking is itself imperfect, even a perfect prediction would not score an SCS of 1.0. The achievable maximum is 0.93 for the agent and 0.90 for objects, well above the scores our models attain.

Dotted extensions are projections beyond 30k hours.

Agent fidelity

Hand SCS ↑
≈0.785
Hand SCS ↑30k hours

Manipulated object fidelity

Object SCS ↑
0.527
Object SCS ↑30k hours
Ego-centric video comes with 3D hand and body keypoints. We can project them onto the image plane as a coloured skeleton — hue for finger identity, value for depth — encode it with the same frozen VAE as the video, and add the tokens directly to the video tokens. The model no longer has to infer the body from action tokens.
Finding 1
Conditioning saturates the agent with far less data.
Supplying the projected skeleton lets 300 hours reach the agent fidelity that unconditioned training attains only at roughly 15k hours — a ≈50× reduction — and both designs converge to nearly the same limit (0.800 and 0.811).
Finding 2
With the agent saturated, object interaction converges far below it.
Object fidelity improves by only 0.09 SCS across two orders of magnitude and decelerates (+0.056 from 300 to 3k hours, +0.035 from 3k to 30k). Its fitted asymptote is 0.565, of which 93% is already realized at 30k hours.
Finding 3
Neither more data nor a larger model closes the gap.
Scaling the model 4× (Super) leaves agent fidelity virtually unchanged and raises the object ceiling by only 0.034 SCS; closing the remaining 0.205 to the agent ceiling at that rate would take roughly six more quadruplings.
03 · Design

Can object interaction modeling be improved?

The previous section established a low limit on object fidelity that additional data is unlikely to overcome. We now ask whether changing the supervision can.

Object-centric adaptive noise scheduling

Combining each point’s 3D displacement and tracking reliability with a soft spatial prior around the projected hands, and expanding the resulting scores over neighbouring latent cells, yields a dynamic region map. Within these regions we raise the noise level, so the denoising target can no longer be satisfied by propagating local appearance and must be predicted from context and dynamics. Reweighting adds a small further improvement in combination at no extra cost, so we include it in our final configuration.

Object fidelity across scales

Object SCS ↑
0.546
Object SCS ↑30k hours
04 · Failure modes

Remaining failure modes.

Our method improves agent fidelity and some aspects of object interaction, as the comparisons above show, but several failure modes remain. Below are examples of each.

Two factors plausibly explain what remains. First, irreducible uncertainty: how an object responds to contact depends on mass, friction and stiffness, none perfectly recoverable from a single frame. Second, the parts that matter most are not observable in pixels at all: contacts and forces are never visible, and the object geometry under the hand is exactly what the hand occludes.

05 · Humanoid transfer

Transfer to humanoids.

We fine-tune Cosmos 3 Nano on ground-truth demonstrations from 11 tasks from the simulated bimanual benchmark of Yang et al. (2025), conditioning on a single start frame and the 138-dimensional humanoid state trajectory and predicting the following 16 frames (17 total at 10 Hz). All evaluations are performed on EgoVLA policy rollouts rather than demonstrations.

Pre-training on human video improves humanoid predictions, and our conditioning and supervision design brings further gains. However, accurately rendering the hands still does not guarantee accurate predictions of their effects on objects.

Humanoid samples

SIMULATOR ROLLOUTS AND SKELETON-CONDITIONED PREDICTIONS
SimulatorPrediction

matchSuccessful can unloading; prediction reproduces success.

SimulatorPrediction

matchUnsuccessful unloading attempt; prediction reproduces failure.

SimulatorPrediction

matchLaptop remains closed; prediction matches.

Where it still misses: object dynamics after contact.

MISSED VERDICTS · SAME FAILURE FAMILY AS HUMAN VIDEO
SimulatorPrediction

missBox stays in place instead of turning with the hand.

SimulatorPrediction

missHand rises, but the laptop stays closed.

SimulatorPrediction

missCan stays high instead of dropping as fingers open.

Quantitative results

Success agreement measures whether generated clips support the same success or failure judgment as the corresponding simulator clips.

Humanoid evaluation by model variant. Success agreement concerns selected 17-frame windows, not whole episodes.
VariantHand SCS ↑Object SCS ↑Success agreement ↑
Cosmos 3 (released)0.630.700.67
+HUMAN (30k hours)0.70+0.070.73+0.030.71+0.04
Ours0.88+0.170.79+0.060.81+0.10

Object fidelity is high in absolute terms for every variant (0.70–0.79), which we attribute to the simulated environment: its objects are rigid and relatively large, so the conditions that make object dynamics hard in ego-centric human video are largely absent. Conclusions about interaction fidelity thus depend heavily on the difficulty of the evaluation setting.