MVG-WAM Read the paper

GEOMETRY-AWARE ROBOT LEARNING

MVG–WAM

One world.
Every view.

Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation

Wenbo Chen1,* Tianfu Li1,* Haoxuan Xu2,* Zhihao Cao3 Zhenghan Chen4 Zhengming Zhu5 Zizhou Luo6 Guosheng Yang1 Yuan Liu2 Lujia Wang1 Wen Chen1 Haoang Li1

Explicit multi-view geometry.
No additional embodied pretraining.

FIELD NOTES / 001COBOT MAGIC
EXTERNAL VIEW3×
Scene camera observes the plate and radish
01 Scene
Left wrist camera observation
02 Left wrist
Right wrist camera observes the radish near the gripper
03 Right wrist
Different perspectives. A shared physical world.

1 The Hong Kong University of Science and Technology (Guangzhou)

2 The Hong Kong University of Science and Technology

3 ETH Zurich

4 Zhejiang University

5 EPFL

6 University of Zurich

* Wenbo Chen, Tianfu Li, and Haoxuan Xu contributed equally to this work.

MEASURED IN ACTION

Geometry that
makes a difference.

99.1%

LIBERO

40 tasks · average success
92.07%

RoboTwin 2.0

50 tasks · average success
91.3%

Real world

137 / 150 successful trials

01 / IN THE REAL WORLD

See the whole scene.
Get the details right.

Grasping, precise alignment, and multi-stage interaction on a dual-arm Cobot Magic. One policy, three tasks.

Pick up the radish and place it on the plate.

3× AS PRESENTED

50 evaluation trials per task. The external recording is separate from the three camera observations supplied to the policy. Videos show selected successful executions.

A CLOSER LOOK

Where interaction succeeds.

Selected successes and failure cases.
10× speed as presented in the source demos.

These are representative rollouts, not matched initial-state tests. The reported overall success rates use the same 150-trial evaluation protocol: MVG-WAM 91.3%, FastWAM-Joint 78.7%, π0.5 89.3%, and MVG-WAM without the epipolar residual 87.3%.

02 / THE GEOMETRY BEHIND THE ACTION

Multiple cameras.
One representation.

Synchronized observations are projections of the same physical world. MVG-WAM makes those geometric relationships explicit inside a world–action model.

p′T F p = 0

Cross-view correspondences
constrained by calibrated geometry

01

Connect the views

Epipolar-constrained retrieval aggregates dense DINOv3 features into a global state, linking geometrically admissible observations across cameras.

02

Keep each camera’s context

Jointly encoded VGGT-Ω registers provide view-indexed geometric states. Camera-aware routing supplies each video region with its matching context.

03

Ground the future in scale

Multi-horizon metric-depth supervision grounds the video representation during training. The depth decoder is removed during action rollout.

The global and per-view geometric states condition the video stream. Action tokens access the geometry-enhanced representation through mixed attention.
AT DEPLOYMENT

RGB + language + proprioception + camera calibration

Action sequenceGeometry is encoded once per observation and reused across denoising steps. No depth decoding during control.
Read the full abstract

World–Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit. This makes it harder to connect global scene context with the local geometry required for interaction. We introduce the Multi-View Geometry-Aware World–Action Model (MVG-WAM), which organizes these observations as related projections of one physical world rather than separate images on a canvas. Our model combines an epipolar-constrained global state with view-indexed geometric states jointly inferred from synchronized observations. Camera-aware routing supplies each video region with its corresponding geometric context and the shared global state, explicitly structuring the representation used for action prediction. We further ground the geometry-aware representation in metric scale through multi-horizon future-depth supervision, without requiring depth decoding during action rollout. MVG-WAM achieves average success rates of 99.1% on LIBERO and 92.07% on RoboTwin 2.0, demonstrating competitive performance across both benchmarks. Real-world experiments on Cobot Magic further demonstrate a 91.3% success rate across 150 trials spanning three manipulation tasks.

03 / INSIDE THE SIMULATOR

More tasks.
More perspectives.

Explore six paired demonstrations from RoboTwin 2.0 and LIBERO. Each pair shows MVG-WAM alongside FastWAM-Joint.

Selected examples illustrate successful MVG-WAM rollouts and baseline failures from the presentation. Quantitative results below summarize the full benchmark evaluation.

THE FULL PICTURE

Performance across benchmarks.

Success rates (%) reported in the paper. Additional embodied pretraining is shown separately.

04 / WHAT THE MODEL LEARNS

Geometry at the
point of contact.

Short-horizon predictions preserve sharper object boundaries and more coherent gripper–object geometry in the paper’s qualitative examples.

Both models use the same current observation and are compared with the same future target.
RGB PREDICTION / PSNR31.79 dB

+3.74 dB over FastWAM-Joint

STRUCTURAL SIMILARITY / SSIM0.914

Compared with 0.842 for FastWAM-Joint

TEACHER-FORCED DEPTH / MAE2.12 cm

Metric information retained in the denoising representation

625 clips from 125 validation trajectories. Depth readout uses noisy future target latents and is a representation diagnostic, not current-observation-only depth forecasting.

05 / EXPLORE THE WORK

A shared world.
A new perspective.

Read the complete method, evaluation protocols, and analysis in the arXiv preprint.

The code will be released in the project repository.

Citation

arXiv:2609.37793 · 2026.

@misc{chen2026mvgwammultipleviewgeometryaware,
  title = {MVG-WAM: Multiple View Geometry-Aware
           World-Action Modeling for Robotic Manipulation},
  author = {Wenbo Chen and Tianfu Li and Haoxuan Xu
            and Zhihao Cao and Zhenghan Chen and Zhengming Zhu
            and Zizhou Luo and Guosheng Yang and Yuan Liu
            and Lujia Wang and Wen Chen and Haoang Li},
  year = {2026},
  eprint = {2609.37793},
  archivePrefix = {arXiv},
  primaryClass = {cs.RO},
  url = {https://arxiv.org/abs/2609.37793}
}

Demo