A benchmark for 3D consistency in world models

ORBIA

Do Generated Videos Form a Consistent 3D World?

The ORBIA Team

Overview

ORBIA evaluates whether generated videos form a consistent 3D world by testing preservation of the input scene, persistence of newly generated content, and agreement across views during exploration and revisit.

6Trajectory protocolsExplore · Revisit · ReturnSee protocols

Evaluating 3D consistency

Each case starts at P0, passes a cross-view query Pq, discovers a new region, revisits it after exploring elsewhere, and returns to the exact starting pose Pr.

Main ORBIA scene: a garden path in a simulated village
Dimension 1Input preservation
Dimension 2Generated-content persistence
Input P0
Reference view

The given image. Its registered depth and anchor mask are the ground truth.

Query Pq
Cross-view check

A new viewpoint on input-visible surfaces, compared with the registered RGB-D and anchors.

First visit Pfirst
New region

A region outside the input view. The first generated observation becomes its reference.

Revisit Previsit
Same region again

After exploring elsewhere, the same region should look and measure the same.

Return Pr
Same pose as P0

Back at the start: the view should match the input image and its depth.

Dimension 3 · across the whole route

3D self-consistency

Views sampled along the route should describe one 3D scene.

View 2
View 1
View 3

Depth reprojected between views agrees on shared surfaces.

View 2
View 1
View 2
View 3

The same surface keeps the same visual features in every view.

View 2
View 1
View 3

Depth scale stays uniform across a 4 × 4 grid of image regions.

Target view
Rendered from the reconstruction

A 3D scene built from 64 context views must render 16 held-out views.

Click a step, card or layer to explore

Exploration and revisit protocols

Every template starts at P0 and returns to the same pose at Pr. They differ in parallax, viewing direction, and how far the camera wanders before coming back.

  • P0 / Pr
  • Pq
  • Pfirst
  • Previsit
  • Camera view
  • Input anchor
  • New object
  • Outward
  • Return
01

Forward–backward

Start · P₀

Push forward past the anchor into new space, then back straight out to P₀.

02

Lateral return

Start · P₀

Truck sideways while the camera keeps facing the anchor, then slide back.

03

Yaw return

Start · P₀

Pan past the first-visit pose, keep exploring, then revisit the same pose on the way back.

04

Arc return

Start · P₀

Sweep around the anchor on an arc, come home on a different arc.

05

Orbit around an anchor

Start · P₀

Circle the anchor while looking at it: visible surfaces and occlusions keep changing.

06

Peek-and-return

Start · P₀

Peek past the anchor into an unseen region, return, then revisit it on a second excursion.

Leaderboard

Results for 27 models across control, visual quality, three dimensions of 3D consistency, and temporal stability.

Model
76.3279.9374.7669.2670.2868.8484.05
74.7781.6471.3768.7970.2561.5182.31
72.7066.4776.4868.8166.4368.0283.18
72.2470.7475.0265.9161.4265.2682.69
71.9575.5167.5575.2368.5860.1678.96
68.0764.1874.0366.9360.8254.6176.41
67.4443.2376.9772.8467.7072.5480.39
67.0161.6871.0470.1756.6258.9276.31
66.7255.2674.4361.0154.4161.2583.15
65.9858.4875.3461.5452.0057.1177.28
65.5045.2672.0666.8469.8470.3477.32
65.2338.5877.3869.1463.3271.3079.30
65.0859.3973.2263.1951.9554.1075.03
64.3460.4968.7858.2346.3159.1778.25
63.7227.9279.0466.0469.7574.5579.73
63.2736.0977.0763.1257.4471.3178.96
62.7535.7074.2566.7866.2970.1074.71
62.7349.1273.8459.5650.7158.9375.37
62.7260.9268.0558.6543.7246.3278.04
62.0032.4878.0963.3655.7969.4277.51
61.7948.9974.9660.9648.7048.7374.82
61.2526.5177.9664.8160.2571.6477.30
60.6439.1670.7063.0860.1763.5372.51
60.5531.3176.4665.9654.3564.1975.79
58.8747.0666.9757.4745.2248.3276.32
57.1823.7677.3058.8947.7764.1174.22
52.3028.4561.3059.3844.7459.9167.31
FirstSecondThird

Ranks are calculated across all 27 models for each metric. Click a model for its profile.
Control interfaces reflect the signals actually received by the model: direct action inputs are labeled Action; actions converted to camera parameters before model input are labeled Camera.

Key findings

Performance varies across evaluation tiers

SolarWM-H3 leads overall, but no model leads every tier.

Visual quality and 3D consistency

Beautiful frames can still describe an inconsistent 3D world.

Camera control and geometric consistency

Following the camera path and maintaining coherent geometry are distinct abilities.

Preserving input and generated content

Recovering the starting scene does not guarantee remembering newly generated content.

Consistency during extended exploration

Extended exploration weakens both memory and camera control.

Model comparisons

Compare models given the same input and requested trajectory, with synchronized playback and registered evaluation events.

How faithfully does generation follow the requested camera motion?

Input image for case_0164
Input scene

A desolate concrete construction site features weathered gray walls and a central pillar, with scattered debris including hollow cinder blocks, broken slabs, and rubble on the dusty ground. The scene is devoid of vegetation, water, or furniture, presenting a stark, industrial atmosphere under diffuse daylight. Materials consist primarily of rough concrete and sand, with visible textures of wear and age on the surfaces. Spatial relations remain consistent, showing a cornered area enclosed by high walls, with no signs of movement or human presence. The environment conveys abandonment and stillness, emphasizing the raw, unfinished state of the structure.

Registered query reference
Query reference
0:000:22
Jump to
Lyra2Case 97.41
Overall Control81.64
EVOKECase 51.74
Overall Control66.47
SolarWM-H3Case 50.37
Overall Control79.93
HY-WorldPlayCase 31.49
Overall Control45.26
Seedance2.5Case 8.11
Overall Control43.23
AstraCase 7.50
Overall Control28.45

Case scores evaluate this rollout. Overall scores are the model’s aggregate on the selected tier. All videos use the same playback speed. The timeline and event timestamps remain in original rollout time.

Dataset construction

Each case combines registered reference views, verified landmarks and a trajectory designed for exploration and revisit.

Citation

@misc{lu_orbia,
  title = {Orbia: Do Generated Videos Form a Consistent 3D World?},
  author = {Dongyue Lu and Guian Fang and Hur Lim and Yonggang Wu and Yanrong Wang and Hongshan Chen and Shengju Qian and Yue Liao and Wei Chow and Lingdong Kong and Tianxin Huang and Yu Yang and Hongji Yang and Junchao Huang and Alan Liang and Yihao Wang and Zihan Wang and Rong Li and Hanlin Chen and Xin Wang and Wei Tsang Ooi and Mike Zheng Shou and Shuicheng Yan},
  note = {ORBIA benchmark manuscript}
}