EMBODIED WORLD MODELS / 2026

CausalWM v1

Causal Chain-of-Thought Reasoning
for Embodied World Model

A 16B embodied world model that makes physical reasoning explicit. CausalWM predicts optical flow, then 3D pointmaps, then future video—using each intermediate prediction as context for the next. Language instructions and robot actions provide complementary ways to guide the imagined future.

O → F → P → VCAUSAL REASONING TRAJECTORY
01 / OPTICAL FLOW
MotionREUSE AS CONTEXT
02 / POINTMAPS
GeometryREUSE AS CONTEXT
03 / RGB VIDEO
FuturePREDICTED OBSERVATION

Flow → Pointmaps → RGB

01 / BENCHMARK · SEP. 11, 2026

TriWorldBench

Evaluating one consistent robot world across head, left-wrist and right-wrist views.

Official leaderboard
#1OF 36 MODELS
CausalWM listed as CWM66.04TWB-SCORE

OFFICIAL LEADERBOARD / TOP FIVE

RankModelTWB-Score ↑
01CausalWMCWM66.04
02dream4act65.66
03BWM65.54
04WoVR_Plus65.39
05PhyxWM64.26

September 11, 2026 leaderboard snapshot.
Top five of all 36 published models. Scores and ranks reproduced from the official leaderboard.

Download leaderboard values

CAUSALWM / INDIVIDUAL METRICS

Seven metrics in the top two.

2 first-place and 5 second-place results across the 19 official evaluation metrics.

Perspective#191.16
Image Quality#143.24
VLM Consistency I#285.83
VLM Consistency III#295.64
VQA Consistency#270.03
Instruction Following#276.78
Subject Consistency#284.55
0Higher is better100

Official individual-metric ranks among all 36 published models · September 11, 2026. Scores are on a 0–100 scale.

Download all 19 metric values

PAI-BENCH / ROBOT DOMAIN

Physical reasoning.
From language to future.

Given one observation and a language instruction, CausalWM generates a future through explicit motion and geometry predictions.

89.9RO SCORE / 100Reaches SOTA

State-of-the-art performance on language-conditioned robot-domain generation.

Direct future prediction compared with CausalWM's optical-flow, pointmap and RGB reasoning chain. CausalWM scores 66.04 on TriWorldBench and 89.9 on the PAI-Bench robot domain. The PAI-Bench figure compares Wan2.2-I2V-A14B, Veo 3, Cosmos3-Super and CausalWM.
Reasoning paradigms and headline results · manuscript illustrationOpen full figure
Evaluation

174 prompts × 5 seeds
870 generations · 913 binary VQA questions

Generation

4 / 4 / 4 denoising steps · CFG 1
121 frames at 640 × 480

Judge

Qwen3-VL-235B-A22B-Instruct
Equal weight per video · 0–100 score

CausalWM and Cosmos3-Super (89.7) are evaluated locally; other baseline scores come from the official leaderboard. CausalWM uses prompts rewritten to match its pretraining captions, while Cosmos3-Super follows its technical report’s inference settings. The manuscript notes that local evaluation does not exactly reproduce the leaderboard’s absolute scores. The few-step study below is a separate denoising-budget comparison.

Download all nine PAI-Bench results

PAI-BENCH / FEW-STEP GENERATION

One step per stage.
Three steps to the future.

88.84RO score / 100
5.16×Speedup vs. 20/20/20

02 / THE FRAMEWORK

Motion. Geometry.
Then, the future.

One shared diffusion Transformer predicts optical flow, then pointmaps, then future RGB. Each completed stream is held fixed as context for the next; causal attention blocks information from later stages.

CausalWM framework: visual context and multimodal conditioning guide causal optical-flow, pointmap and future-RGB prediction, with three training stages and in-context control.
The CausalWM framework · shared backbone, causal reasoning and visual controlOpen full figure

03 / CASE STUDIES

See the steps
between now and next.

Explore generated motion, geometry and future video across three language-conditioned robot tasks.

LANGUAGE INSTRUCTION

The robotic gripper picks up the blue bottle on the bathroom counter and places it in the open drawer.

121 frames · 16 fps · 640 × 480 · 10/10/10 denoising steps
01 / Optical flow02 / Pointmaps03 / Future RGB
The three predicted streams at matching video frames. Use the player to pause, scrub or replay.

THE INFORMATION FLOW

A causal order.
An explicit dependency.

Each stream can read the observation and itself or earlier streams. Later-to-earlier paths are blocked; attention within each stream remains bidirectional.

The matrix shows stream visibility, not measured attention weights.

KEYS / CONTEXT →
OFPV
O
F
P
V
AllowedBlockedO → F → P → V

ACTION-CONDITIONED / TRIWORLDBENCH

One action.
Three viewpoints.

Robot trajectories are rendered as URDF control videos to guide synchronized head and wrist views.

TASK / EPISODE 25

Pick up the medium roller with narrower tips using both arms equally

Selected best-of-8 submission · 96 frames · 30 fps
Head cameraLeft wristRight wrist
URDF CONTROL INPUTGENERATED RGB
Top: the black-background URDF control supplied to the model. Bottom: the generated views at the same frames. This action-conditioned example uses rendered robot control; the flow and pointmap predictions are shown in the language-conditioned cases above.

04 / THE DATA FOUNDATION

Many embodiments.
A shared physical world.

Human activity, real robots and simulation contribute about 31K source hours, curated into 20K hours of interaction experience.

20Source families
31,230.8Source hours · ≈31K
19,916.2Retained hours · ≈20K
63.8%Hours retained

Explore the composition.

Outer ring: datasets · Inner ring: embodiment groups

Human hands: 8,206.9 hoursSingle-arm: 2,335.3 hoursDual-arm: 3,359.2 hoursSingle- and dual-arm: 5,662.7 hoursHuman hands + robot arms: 352.2 hoursEgo4D · 798.5 hEgocentric-10K · 6,878.2 hEgoDex · 513.4 hEPIC-Kitchens · 16.4 hH2O · 0.4 hBridgeData V2 · 59.0 hDROID · 189.4 hOXE · 336.5 hRoboCasa365 · 1,502.9 hRT-1 · 247.5 hAgiBot-World 2026 · 220.9 hAgiBot-World Beta · 2,022.1 hGalaxea · 485.9 hHumanoid-Everyday · 28.5 hRoboCOIN · 488.0 hRoboTwin 2.0 · 113.8 hInternData-A1 · 1,507.6 hRoboMIND · 150.5 hRoVid-X · 4,004.6 hEgoVerse · 352.2 hRETAINED HOURS6,878.234.5% of the pool02 / 20 SOURCES
Hover to explore · click to select

Human hands

Egocentric-10K

Source hours9,980.3
Retained hours6,878.2
Retention68.9%
Camera views
Egocentric
Task coverage
Industrial operations
Human hands41.2%
Single-arm11.7%
Dual-arm16.9%
Single- and dual-arm28.4%
Human hands + robot arms1.8%

Synchronized camera views of the same trajectory are counted once. Totals follow the manuscript’s data table. Chart shares use the displayed rows, which sum to 31,230.9 source hours and 19,916.3 retained hours; each differs from the reported total by 0.1 hour.

01 / HUMAN EGOCENTRIC

Everyday interaction.

Daily activity, tool use and dexterous hand–object interaction across human viewpoints.

Ego4D · Egocentric-10K · EgoDex · EPIC-Kitchens · H2O · EgoVerse
02 / REAL ROBOTS

Diverse embodiments.

Single-arm, bimanual, mobile-manipulation and humanoid platforms with varied cameras and end effectors.

AgiBot · Galaxea · RoboCOIN · RoboMIND · DROID · RT-1 · BridgeData V2 · OXE · Humanoid-Everyday · RoVid-X
03 / SIMULATION

Broader physical coverage.

Additional objects, trajectories, scenes and embodiments with explicit control and annotation.

InternData-A1 · RoboCasa365 · RoboTwin 2.0 · AgiBot-World 2026 digital twins

FROM SOURCE VIDEOS TO TRAINING CLIPS

Curate the interaction.
Clarify the instruction.

Processing routes adapt to each source. Clean recordings may bypass individual filters; suitable native annotations are normalized and reused.

  1. 01

    Validate duration

    Check frame rates, timestamps and usable clip length.

  2. 02

    Screen motion

    Apply source-specific flow thresholds for excessive or insufficient motion.

  3. 03

    Check the action

    Keep recognizable actors and task-relevant physical interactions.

  4. 04

    Segment & recaption

    Build contiguous events with concise, action-focused instructions.