
Input P0The given image. Its registered depth and anchor mask are the ground truth.
Do Generated Videos Form a Consistent 3D World?
The ORBIA Team
ORBIA evaluates whether generated videos form a consistent 3D world by testing preservation of the input scene, persistence of newly generated content, and agreement across views during exploration and revisit.
Each case starts at P0, passes a cross-view query Pq, discovers a new region, revisits it after exploring elsewhere, and returns to the exact starting pose Pr.


Input P0The given image. Its registered depth and anchor mask are the ground truth.

Query PqA new viewpoint on input-visible surfaces, compared with the registered RGB-D and anchors.

First visit PfirstA region outside the input view. The first generated observation becomes its reference.

Revisit PrevisitAfter exploring elsewhere, the same region should look and measure the same.

Return PrBack at the start: the view should match the input image and its depth.
Views sampled along the route should describe one 3D scene.
View 2
View 1
View 3Depth reprojected between views agrees on shared surfaces.
View 2


The same surface keeps the same visual features in every view.



Depth scale stays uniform across a 4 × 4 grid of image regions.
Target view
Rendered from the reconstructionA 3D scene built from 64 context views must render 16 held-out views.
Every template starts at P0 and returns to the same pose at Pr. They differ in parallax, viewing direction, and how far the camera wanders before coming back.
Push forward past the anchor into new space, then back straight out to P₀.
Truck sideways while the camera keeps facing the anchor, then slide back.
Pan past the first-visit pose, keep exploring, then revisit the same pose on the way back.
Sweep around the anchor on an arc, come home on a different arc.
Circle the anchor while looking at it: visible surfaces and occlusions keep changing.
Peek past the anchor into an unseen region, return, then revisit it on a second excursion.
Results for 27 models across control, visual quality, three dimensions of 3D consistency, and temporal stability.
| Model | |||||||
|---|---|---|---|---|---|---|---|
| 76.32 | 79.93 | 74.76 | 69.26 | 70.28 | 68.84 | 84.05 | |
| 74.77 | 81.64 | 71.37 | 68.79 | 70.25 | 61.51 | 82.31 | |
| 72.70 | 66.47 | 76.48 | 68.81 | 66.43 | 68.02 | 83.18 | |
| 72.24 | 70.74 | 75.02 | 65.91 | 61.42 | 65.26 | 82.69 | |
| 71.95 | 75.51 | 67.55 | 75.23 | 68.58 | 60.16 | 78.96 | |
| 68.07 | 64.18 | 74.03 | 66.93 | 60.82 | 54.61 | 76.41 | |
| 67.44 | 43.23 | 76.97 | 72.84 | 67.70 | 72.54 | 80.39 | |
| 67.01 | 61.68 | 71.04 | 70.17 | 56.62 | 58.92 | 76.31 | |
| 66.72 | 55.26 | 74.43 | 61.01 | 54.41 | 61.25 | 83.15 | |
| 65.98 | 58.48 | 75.34 | 61.54 | 52.00 | 57.11 | 77.28 | |
| 65.50 | 45.26 | 72.06 | 66.84 | 69.84 | 70.34 | 77.32 | |
| 65.23 | 38.58 | 77.38 | 69.14 | 63.32 | 71.30 | 79.30 | |
| 65.08 | 59.39 | 73.22 | 63.19 | 51.95 | 54.10 | 75.03 | |
| 64.34 | 60.49 | 68.78 | 58.23 | 46.31 | 59.17 | 78.25 | |
| 63.72 | 27.92 | 79.04 | 66.04 | 69.75 | 74.55 | 79.73 | |
| 63.27 | 36.09 | 77.07 | 63.12 | 57.44 | 71.31 | 78.96 | |
| 62.75 | 35.70 | 74.25 | 66.78 | 66.29 | 70.10 | 74.71 | |
| 62.73 | 49.12 | 73.84 | 59.56 | 50.71 | 58.93 | 75.37 | |
| 62.72 | 60.92 | 68.05 | 58.65 | 43.72 | 46.32 | 78.04 | |
| 62.00 | 32.48 | 78.09 | 63.36 | 55.79 | 69.42 | 77.51 | |
| 61.79 | 48.99 | 74.96 | 60.96 | 48.70 | 48.73 | 74.82 | |
| 61.25 | 26.51 | 77.96 | 64.81 | 60.25 | 71.64 | 77.30 | |
| 60.64 | 39.16 | 70.70 | 63.08 | 60.17 | 63.53 | 72.51 | |
| 60.55 | 31.31 | 76.46 | 65.96 | 54.35 | 64.19 | 75.79 | |
| 58.87 | 47.06 | 66.97 | 57.47 | 45.22 | 48.32 | 76.32 | |
| 57.18 | 23.76 | 77.30 | 58.89 | 47.77 | 64.11 | 74.22 | |
| 52.30 | 28.45 | 61.30 | 59.38 | 44.74 | 59.91 | 67.31 |
Ranks are calculated across all 27 models for each metric. Click a model for its profile.
Control interfaces reflect the signals actually received by the model: direct action inputs are labeled Action; actions converted to camera parameters before model input are labeled Camera.
SolarWM-H3 leads overall, but no model leads every tier.
Beautiful frames can still describe an inconsistent 3D world.
Following the camera path and maintaining coherent geometry are distinct abilities.
Recovering the starting scene does not guarantee remembering newly generated content.
Extended exploration weakens both memory and camera control.
Compare models given the same input and requested trajectory, with synchronized playback and registered evaluation events.
How faithfully does generation follow the requested camera motion?

A desolate concrete construction site features weathered gray walls and a central pillar, with scattered debris including hollow cinder blocks, broken slabs, and rubble on the dusty ground. The scene is devoid of vegetation, water, or furniture, presenting a stark, industrial atmosphere under diffuse daylight. Materials consist primarily of rough concrete and sand, with visible textures of wear and age on the surfaces. Spatial relations remain consistent, showing a cornered area enclosed by high walls, with no signs of movement or human presence. The environment conveys abandonment and stillness, emphasizing the raw, unfinished state of the structure.

Case scores evaluate this rollout. Overall scores are the model’s aggregate on the selected tier. All videos use the same playback speed. The timeline and event timestamps remain in original rollout time.
Each case combines registered reference views, verified landmarks and a trajectory designed for exploration and revisit.







Open a scene to inspect its input and query RGB images and registered depth.
@misc{lu_orbia,
title = {Orbia: Do Generated Videos Form a Consistent 3D World?},
author = {Dongyue Lu and Guian Fang and Hur Lim and Yonggang Wu and Yanrong Wang and Hongshan Chen and Shengju Qian and Yue Liao and Wei Chow and Lingdong Kong and Tianxin Huang and Yu Yang and Hongji Yang and Junchao Huang and Alan Liang and Yihao Wang and Zihan Wang and Rong Li and Hanlin Chen and Xin Wang and Wei Tsang Ooi and Mike Zheng Shou and Shuicheng Yan},
note = {ORBIA benchmark manuscript}
}