Video2World

Benchmarking Coding Agents for Interactive World Modeling from Embodied Videos

Jinzhou Tang1,2*, Zijun Zhang1*, Jing Yang1, Yuchen Yan1, Kun Zhou1†, Lingjun Mao1, Ruobing Han1, Jinglin Cao1, Wenpeng Xu1, Lukun He1, Minghao Fu1,2, Fan Feng1, Biwei Huang1,2

1Aether AI2University of California, San Diego

*Equal Contribution†Corresponding Author and Project Lead

PaperCodeDatasetLeaderboardComing soonOne-click EvalSoon

From one embodied video
to an interactive world,
via coding agents.

Explore the benchmark

Still challenging.

On V2WScore, the best agent reaches 48.5 out of 100. The human-assisted reference reaches 74.2.

Explore the tasks See the results

Video2World

Benchmarking Coding Agents for Interactive World Modeling from Embodied Videos

Jinzhou Tang1,2*, Zijun Zhang1*, Jing Yang1, Yuchen Yan1, Kun Zhou1†, Lingjun Mao1, Ruobing Han1, Jinglin Cao1, Wenpeng Xu1, Lukun He1, Minghao Fu1,2, Fan Feng1, Biwei Huang1,2

1Aether AI2University of California, San Diego

*Equal Contribution†Corresponding Author and Project Lead

PaperCodeHugging FaceLeaderboardcoming soon
Video2World teaser (paper Figure 1): real robot and human videos are turned into interactive simulated worlds by off-the-shelf coding agents; a dataset composition wheel of 222 tasks and a chart of model performance on functionality, geometry and dynamics.
Figure 1Video2World benchmarks the conversion of embodied videos into simulated worlds. Given an RGB video and simulator tools, a coding agent constructs the environment and robot behavior, then revises its implementation using execution feedback. Evaluation examines geometric fidelity, physically generated object motion, task completion.

00 / Abstract

Looks right, still fails.

Building a simulator from a real video still takes hand engineering. Video2World asks whether coding agents can do it on their own. An agent watches one robot or human video and writes the scene, its physics and the robot's actions as code. We run that code and check three things: is the geometry right, does the motion match the video, and does the robot finish the task. Across 222 tasks from 189 videos, the best of nine frontier agents finishes 25.5% of tasks, while a human-assisted reference reconstruction reaches 58.8%. And accuracy alone is not enough: GPT-6 Astra builds the most accurate shapes, yet finishes only 10.7%.

Video2World in 90 seconds: one video, nine coding agents, and why an accurate world can still fail.

01 / Tasks

Explore the tasks.

222 instances in 39 task families, from robot, egocentric and third-person human videos.

Task families
39
Instances
222
Video clips
189
Category I · 12 families · 68 instances

Assemble & insert

Seat, insert and assemble parts, where contact geometry decides whether the task can succeed.

Representative · F07 · square-table leg 1 (FurnitureBench)FurnitureBench · SAPIEN
Show detailsHide details

Task families

  • F01Bulb seating8
  • F02Lampshade placement8
  • F03Cabinet-top assembly6
  • F04Cabinet-door assembly5
  • F05Chair-back assembly1
  • F06Drawer insertion5
  • F07Corner-leg insertion26
  • F08Chair-nut seating3
  • F09Central-base assembly2
  • F30Held rigid-object insertion2
  • F32Held-wallet insertion1
  • F36Air-fryer basket extraction1

Where the instances come from

DatasetRobotSimulatorViewpointInstances
FurnitureBenchFranka PandaSAPIENThird-person robot video64
In-house householdG1 + OmnipickerIsaac SimRobot head camera4
Category II · 13 families · 51 instances

Pick & place

Robot demonstrations of picking, lifting, holding and placing objects, from a table-top arm to a humanoid at home.

Representative · F10 · cup into bowl (DROID)DROID · SAPIEN
Show detailsHide details

Task families

  • F10DROID rigid manipulation18
  • F22Book placement1
  • F23Box placement1
  • F24Bottle pickup1
  • F25Cracker-box pickup1
  • F26Remote-control pickup1
  • F27Lift and hold · episode poses6
  • F28Held-object placement · episode poses2
  • F29Lift and hold · reviewed poses2
  • F31Wallet lifting1
  • F33Lift and hold · object-origin points5
  • F34Lift and hold · visible-surface points9
  • F35Held placement · visible-surface points3

Where the instances come from

DatasetRobotSimulatorViewpointInstances
In-house householdG1 + OmnipickerIsaac SimRobot head camera33
DROIDPanda + Robotiq 2F-85SAPIENThird-person robot video18
Category III · 8 families · 66 instances

Human video

Human hand demonstrations, egocentric or third-person, re-enacted by a robot arm or by paired dexterous hands with the same object-level goal.

Representative · F16 · bowl placement (HOI4D)HOI4D · SAPIEN
Show detailsHide details

Task families

  • F12HOT3D placement · arm7
  • F13HOT3D placement · hands7
  • F14DexYCB lifting · arm11
  • F15DexYCB lifting · hands11
  • F16HOI4D placement · arm9
  • F17HOI4D placement · hands9
  • F18OakInk2 placement · arm6
  • F19OakInk2 placement · hands6

Where the instances come from

DatasetRobotSimulatorViewpointInstances
DexYCBFranka PandaSAPIENThird-person human video11
DexYCBPaired Wuji handsSAPIENThird-person human video11
HOI4DFranka PandaSAPIENEgocentric human video9
HOI4DPaired Wuji handsSAPIENEgocentric human video9
HOT3DFranka PandaSAPIENEgocentric human video7
HOT3DPaired Wuji handsSAPIENEgocentric human video7
OakInk2Franka PandaSAPIENThird-person human video6
OakInk2Paired Wuji handsSAPIENThird-person human video6
Category IV · 2 families · 25 instances

Cross-embodiment

One RoboDojo demonstration, rebuilt for two different robots: the same task executed by dual X5 arms and by dual xArm7 arms.

Representative · F20 · stack blocks (RoboDojo, dual X5)RoboDojo · Isaac Sim
Show detailsHide details

Task families

  • F20RoboDojo interaction · dual X513
  • F21RoboDojo interaction · dual xArm712

Where the instances come from

DatasetRobotSimulatorViewpointInstances
RoboDojoDual X5Isaac SimSimulated demonstration13
RoboDojoDual xArm7Isaac SimSimulated demonstration12
Category V · 4 families · 12 instances

Physics-rich

Cloth, rope, soft toys and planar pushing, where the reconstructed world has to respond physically.

Representative · F38 · rope routing (reconstructed twin)Reconstructed twins · MuJoCo
Show detailsHide details

Task families

  • F11DROID cloth folding2
  • F37Planar T-body pushing3
  • F38Rope routing5
  • F39Soft-toy packing2

Where the instances come from

DatasetRobotSimulatorViewpointInstances
Reconstructed twinsxArm7MuJoCoReal video7
Reconstructed twinsxArm7 + pusherSAPIENReal video3
DROIDPanda + cloth gripperSAPIENThird-person robot video2

02 / Evaluation

Build it. Run it. Evaluate it.

IGeometry

Geometric fidelity

Does the rebuilt scene match the one in the video?

Geometry · What is compared

The scene before anything moves

The initial simulated scene is compared with annotated reference surfaces of the source environment, in a common frame. Supports such as the table are kept apart from the objects that are manipulated or receive them.

Geometry · Design logic

Weighted by what the task touches

Errors on manipulated objects, receivers and nearby structures count more than distant clutter. Scene-level correspondence and per-object shape and size are reported separately, so a plausible layout cannot hide a wrong object.

Geometry · Metrics

Three distances, in centimetres

Scene CD
Interaction-weighted Chamfer distance between scene surfaces after bounded registration.
Shape CD
Object shape after rigid alignment, without rescaling.
Size error
Difference in physical object dimensions.

Lower is better.

IIDynamics

Dynamic fidelity

When it runs, does the object move as it did in the video?

Dynamics · What is compared

Object motion, not robot motion

Each method executes its own behaviour in its own scene. The resulting object trajectories are compared with the demonstrated object motion after temporal correspondence, so different control strategies can realise the same interaction.

Dynamics · Design logic

Relative to where each motion starts

Every trajectory is measured from its own initial position, which removes global offsets but keeps motion magnitude and timing. Known object symmetries are respected, deformable objects use surface correspondence, and errors are capped.

Dynamics · Metrics

Where, how it turns, and step by step

Translation APE
Displacement error over the whole motion (cm).
Rotation APE
Orientation error, symmetry-aware (degrees).
Translation RPE
Error in short displacement increments of about 0.2 s (cm).

Lower is better.

IIIFunctionality

Functional correctness

Can the task actually be done in this world?

Functionality · What is compared

The interaction the clip defines

Each task family has object-centric requirements: the initial relations, the essential interaction events and their order, and the terminal goal with its required persistence. Any embodiment, trajectory or timing is allowed if the interaction physically happens.

Functionality · Design logic

It must build before it can count

The build check gates everything: a submission that does not execute scores zero. Milestones along the interaction give partial credit, showing how far an execution gets even when the final goal is missed.

Functionality · Metrics

Success, progress, and the gate

Task Success
All conditions and the terminal goal are satisfied (%).
Task Progress
Share of interaction milestones reached (0–1).
Build Rate
Submissions that produce an executable world (%).

Higher is better.

03 / Benchmark Results

Who builds working worlds?

Video2World benchmark results. Arrows mark whether higher or lower is better.
OverallBuildFunctionalityGeometryDynamics
#Method
—Human-assisted Ref reference74.24100.058.80.702.190.150.439.4041.092.42
1Claude Fable 5.148.5293.125.50.535.351.002.7818.71104.618.25
2GPT-6 Astra43.5593.110.70.336.230.691.6220.7895.086.19
3Claude Opus 541.5692.816.00.346.011.143.1319.37110.858.42
4Kimi K333.1992.04.80.218.081.414.2327.28111.1614.29
5GPT-5.6 Sol31.0092.52.50.117.791.373.7631.68105.0318.00
6DeepSeek-V4.1 Flash30.5492.44.10.138.591.504.1334.64104.5220.85
7Gemini-3.8 Flash30.1193.12.00.098.571.253.6737.97114.8228.65
8Qwen-3.8 Max-090226.5587.61.50.118.361.514.1341.28110.0627.38
9GLM-5.3 Flash25.6987.72.00.099.581.625.0441.16103.8330.16
SubmitOne-click EvalSoon

04 / Experiments

What the benchmark reveals.

One video, nine worlds.

Source videoPut the tape on the plate. Robot video, DROID.
ReferenceHuman-assisted reference✓ Success
Success
✓ Yes
Progress
1.00
Final distance
2.5 cm to goal

Tape on the plate, 2.5 cm from the goal.

  1. 1Claude Opus 5✓ SuccessOn the plate, 1.9 cm from the goal.
  2. 2GPT-6 Astra✓ SuccessOn the plate, 2.5 cm from the goal.
  3. 3Claude Fable 5.1✓ SuccessOn the plate, 3.0 cm off, despite the roughest scene of the nine.
  4. 4Gemini-3.8 Flash× Progress 0.75Released 6.8 cm from the goal: just outside the 5 cm tolerance.
  5. 5GPT-5.6 Sol× Progress 0.25Grasps, but never reaches the plate.
  6. 6DeepSeek-V4.1 Flash× FailedNever grasps the tape.
  7. 7Qwen-3.8 Max-0902× FailedNever grasps the tape.
  8. 8Kimi K3× FailedNever grasps the tape.
  9. 9GLM-5.3 Flash× FailedNever grasps the tape.

3 of 9 succeed. Claude Fable 5.1 succeeds despite the roughest scene; Gemini-3.8 Flash misses the tolerance by 1.8 cm.

Almost is not done.

The evaluator checks the interaction itself, not how close the end state looks.

1.1 cm from the goal, still a failure.
Failed · progress 0.75

Never passed through the sorting hole.

Wrong pose, task done.
Success

The stack succeeds despite a very different object pose.

28 cm off, still in the bowl.
Success

The package still reaches the target bowl.

Same video, two robots.
Fails on X5 · succeeds on xArm7

The interaction works with one embodiment, but not the other.

What a world costs.

20304050$0.1$0.3$1$3$10Mean API cost per episode (USD) →V2WScoreClaude Fable 5.1*Claude Fable 5.1: $14.03 per episode, 38 min, V2WScore 48.52, success 25.5%GPT-6 Astra*GPT-6 Astra: $6.40 per episode, 15 min, V2WScore 43.55, success 10.7%Claude Opus 5Claude Opus 5: $13.83 per episode, 33 min, V2WScore 41.56, success 16%Kimi K3Kimi K3: $5.51 per episode, 82 min, V2WScore 33.19, success 4.8%GPT-5.6 SolGPT-5.6 Sol: $1.39 per episode, 4 min, V2WScore 31, success 2.5%DeepSeek-V4.1 FlashDeepSeek-V4.1 Flash: $0.23 per episode, 14 min, V2WScore 30.54, success 4.1%Gemini-3.8 FlashGemini-3.8 Flash: $1.53 per episode, 19 min, V2WScore 30.11, success 2%Qwen-3.8 Max-0902Qwen-3.8 Max-0902: $0.67 per episode, 26 min, V2WScore 26.55, success 1.5%GLM-5.3 FlashGLM-5.3 Flash: $0.31 per episode, 53 min, V2WScore 25.69, success 2%
Agent$ / episodemin / episodeV2WScore
Claude Fable 5.1*14.033848.5
GPT-6 Astra*6.401543.5
Claude Opus 513.833341.6
Kimi K35.518233.2
GPT-5.6 Sol1.39431.0
DeepSeek-V4.1 Flash0.231430.5
Gemini-3.8 Flash1.531930.1
Qwen-3.8 Max-09020.672626.6
GLM-5.3 Flash0.315325.7

* Estimated at list price from token counts.

Worlds that keep working.

AgentReuse the old behaviorLet the agent adapt
GPT-6 Astraall 10 misses recovered170/180 10 missed180/180 0 missed
Claude Fable 5.1both misses recovered178/180 2 missed180/180 0 missed
Gemini-3.8 Flash1 of 3 recovered, 2 successes broken177/180 3 missed176/180 4 missed

Move the object or its target 2–20 cm in a rebuilt world and run it again, first with the original robot actions, then after the agent adapts them.

Successes / 180 moved layouts

Gemini-3.8 Flash · tray moved 10 cm

× ReuseBottle left on the table.

✓ AdaptedBottle stands on the moved tray.

GPT-6 Astra · bottle moved 20 cm

× ReuseMisses the tray.

✓ AdaptedPlaced upright on the tray.

Cite Video2World
Download citation
@misc{tang2026video2world,
  title         = {Video2World: Benchmarking Coding Agents for Interactive World Modeling from Embodied Videos},
  author        = {Jinzhou Tang and Zijun Zhang and Jing Yang and Yuchen Yan and Kun Zhou and Lingjun Mao and Ruobing Han and Jinglin Cao and Wenpeng Xu and Lukun He and Minghao Fu and Fan Feng and Biwei Huang},
  year          = {2026},
  eprint        = {2610.04432},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2610.04432}
}