Everyday interaction.
Daily activity, tool use and dexterous hand–object interaction across human viewpoints.
Ego4D · Egocentric-10K · EgoDex · EPIC-Kitchens · H2O · EgoVerseEMBODIED WORLD MODELS / 2026
Causal Chain-of-Thought Reasoning
for Embodied World Model
A 16B embodied world model that makes physical reasoning explicit. CausalWM predicts optical flow, then 3D pointmaps, then future video—using each intermediate prediction as context for the next. Language instructions and robot actions provide complementary ways to guide the imagined future.



Flow → Pointmaps → RGB
01 / BENCHMARK · SEP. 11, 2026
Evaluating one consistent robot world across head, left-wrist and right-wrist views.
Official leaderboardOFFICIAL LEADERBOARD / TOP FIVE
| Rank | Model | TWB-Score ↑ |
|---|---|---|
| 01 | CausalWMCWM | 66.04 |
| 02 | dream4act | 65.66 |
| 03 | BWM | 65.54 |
| 04 | WoVR_Plus | 65.39 |
| 05 | PhyxWM | 64.26 |
September 11, 2026 leaderboard snapshot.
Top five of all 36 published models. Scores and ranks reproduced from the official leaderboard.
CAUSALWM / INDIVIDUAL METRICS
2 first-place and 5 second-place results across the 19 official evaluation metrics.
Official individual-metric ranks among all 36 published models · September 11, 2026. Scores are on a 0–100 scale.
Download all 19 metric valuesPAI-BENCH / ROBOT DOMAIN
Given one observation and a language instruction, CausalWM generates a future through explicit motion and geometry predictions.
State-of-the-art performance on language-conditioned robot-domain generation.

174 prompts × 5 seeds
870 generations · 913 binary VQA questions
4 / 4 / 4 denoising steps · CFG 1
121 frames at 640 × 480
Qwen3-VL-235B-A22B-Instruct
Equal weight per video · 0–100 score
CausalWM and Cosmos3-Super (89.7) are evaluated locally; other baseline scores come from the official leaderboard. CausalWM uses prompts rewritten to match its pretraining captions, while Cosmos3-Super follows its technical report’s inference settings. The manuscript notes that local evaluation does not exactly reproduce the leaderboard’s absolute scores. The few-step study below is a separate denoising-budget comparison.
Download all nine PAI-Bench resultsPAI-BENCH / FEW-STEP GENERATION
02 / THE FRAMEWORK
One shared diffusion Transformer predicts optical flow, then pointmaps, then future RGB. Each completed stream is held fixed as context for the next; causal attention blocks information from later stages.

03 / CASE STUDIES
Explore generated motion, geometry and future video across three language-conditioned robot tasks.
LANGUAGE INSTRUCTION
“The robotic gripper picks up the blue bottle on the bathroom counter and places it in the open drawer.”
121 frames · 16 fps · 640 × 480 · 10/10/10 denoising stepsTHE INFORMATION FLOW
Each stream can read the observation and itself or earlier streams. Later-to-earlier paths are blocked; attention within each stream remains bidirectional.
The matrix shows stream visibility, not measured attention weights.
ACTION-CONDITIONED / TRIWORLDBENCH
Robot trajectories are rendered as URDF control videos to guide synchronized head and wrist views.
TASK / EPISODE 25
“Pick up the medium roller with narrower tips using both arms equally”
Selected best-of-8 submission · 96 frames · 30 fps04 / THE DATA FOUNDATION
Human activity, real robots and simulation contribute about 31K source hours, curated into 20K hours of interaction experience.
Human hands
Synchronized camera views of the same trajectory are counted once. Totals follow the manuscript’s data table. Chart shares use the displayed rows, which sum to 31,230.9 source hours and 19,916.3 retained hours; each differs from the reported total by 0.1 hour.
Daily activity, tool use and dexterous hand–object interaction across human viewpoints.
Ego4D · Egocentric-10K · EgoDex · EPIC-Kitchens · H2O · EgoVerseSingle-arm, bimanual, mobile-manipulation and humanoid platforms with varied cameras and end effectors.
AgiBot · Galaxea · RoboCOIN · RoboMIND · DROID · RT-1 · BridgeData V2 · OXE · Humanoid-Everyday · RoVid-XAdditional objects, trajectories, scenes and embodiments with explicit control and annotation.
InternData-A1 · RoboCasa365 · RoboTwin 2.0 · AgiBot-World 2026 digital twinsFROM SOURCE VIDEOS TO TRAINING CLIPS
Processing routes adapt to each source. Clean recordings may bypass individual filters; suitable native annotations are normalized and reused.
Check frame rates, timestamps and usable clip length.
Apply source-specific flow thresholds for excessive or insufficient motion.
Keep recognizable actors and task-relevant physical interactions.
Build contiguous events with concise, action-focused instructions.