What 30,000 hours of ego-centric video does not teach
Agent and object fidelity
Trained on 30k hours
OverlapGT onlyPredicted onlyEgo-centric data scaling reaches high agent fidelity, but the effect of the agent's actions on the world lags far behind.
The hands land almost exactly where they belong — the paper they are folding does not.
30,012 hours of first-person video.
We train all models on a large-scale ego-centric dataset of everyday human manipulation: 1.15M first-person clips totalling 30k hours (3.24 × 109 frames at 30 fps, mean clip length 94 s), each recording a manipulation activity from a head-mounted camera.
We thank Mecka for their support with the dataset.
More data gets the hands right, but not the objects they handle.
Our models see one first-person frame and the motion of the person's arms and hands, then generate the next 16 frames (1.6 s). We score the agent (the person's hands) and the objects they handle separately, on clips recorded with different people, places and cameras than training. To see what more data would bring, we fit a scaling curve to each result and extend it beyond 30k hours (dotted lines). We then use skeleton conditioning as an instrument for locating the object limit. It additionally supplies these poses as a skeleton projected onto the image, so the model no longer has to infer where the hands appear in the frame, and the hands reach their limit with far less data.
Metric · SCS (structural consistency)
SCS segments and tracks key scene elements through the predicted and ground-truth videos and averages mask IoU, capturing whether the provided actions are accurately reflected in the generated world.
How SCS is computed Example
Because mask tracking is itself imperfect, even a perfect prediction would not score an SCS of 1.0. The achievable maximum is 0.93 for the agent and 0.90 for objects, well above the scores our models attain.
Agent fidelity
Hand SCS ↑Manipulated object fidelity
Object SCS ↑Agent & object appearance LPIPS
Agent appearance
GT-mask LPIPS ↓Manipulated object appearance
GT-mask LPIPS ↓With skeleton conditioning, hand fidelity stays high across 300h–30k hours, while object fidelity remains lower and shows little improvement from 3k to 30k hours.
Can object interaction modeling be improved?
The previous section established a low limit on object fidelity that additional data is unlikely to overcome. We now ask whether changing the supervision can.
Object-centric adaptive noise scheduling
Combining each point’s 3D displacement and tracking reliability with a soft spatial prior around the projected hands, and expanding the resulting scores over neighbouring latent cells, yields a dynamic region map. Within these regions we raise the noise level, so the denoising target can no longer be satisfied by propagating local appearance and must be predicted from context and dynamics. Reweighting adds a small further improvement in combination at no extra cost, so we include it in our final configuration.
Object fidelity across scales
Object SCS ↑Improved agent fidelity
Ground truth.
Cosmos 3 fails to capture the right hand’s motion.
Ours captures the right hand’s motion correctly.
Improved object permanence
Ground truth.
Cosmos 3 baseline drops the tape's identity across the interaction.
Ours preserves it.
Improved interaction modeling
Ground truth.
The baseline flattens the paper mid-interaction.
Ours retains the correct configuration.
Remaining failure modes.
Our method improves agent fidelity and some aspects of object interaction, as the comparisons above show, but several failure modes remain. Below are examples of each.
Two factors plausibly explain what remains. First, irreducible uncertainty: how an object responds to contact depends on mass, friction and stiffness, none perfectly recoverable from a single frame. Second, the parts that matter most are not observable in pixels at all: contacts and forces are never visible, and the object geometry under the hand is exactly what the hand occludes.
Transfer to humanoids.
We fine-tune Cosmos 3 Nano on ground-truth demonstrations from 11 tasks from the simulated bimanual benchmark of Yang et al. (2025), conditioning on a single start frame and the 138-dimensional humanoid state trajectory and predicting the following 16 frames (17 total at 10 Hz). All evaluations are performed on EgoVLA policy rollouts rather than demonstrations.
Pre-training on human video improves humanoid predictions, and our conditioning and supervision design brings further gains. However, accurately rendering the hands still does not guarantee accurate predictions of their effects on objects.
Humanoid samples
matchSuccessful can unloading; prediction reproduces success.
matchUnsuccessful unloading attempt; prediction reproduces failure.
matchLaptop remains closed; prediction matches.
Where it still misses: object dynamics after contact.
missBox stays in place instead of turning with the hand.
missHand rises, but the laptop stays closed.
missCan stays high instead of dropping as fingers open.
Quantitative results
Success agreement measures whether generated clips support the same success or failure judgment as the corresponding simulator clips.
| Variant | Hand SCS ↑ | Object SCS ↑ | Success agreement ↑ |
|---|---|---|---|
| Cosmos 3 (released) | 0.63 | 0.70 | 0.67 |
| +HUMAN (30k hours) | 0.70+0.07 | 0.73+0.03 | 0.71+0.04 |
| Ours | 0.88+0.17 | 0.79+0.06 | 0.81+0.10 |
Object fidelity is high in absolute terms for every variant (0.70–0.79), which we attribute to the simulated environment: its objects are rigid and relatively large, so the conditions that make object dynamics hard in ego-centric human video are largely absent. Conclusions about interaction fidelity thus depend heavily on the difficulty of the evaluation setting.