ACG-WAM

ACG-WAM: World Action Modelingvia Action Conditioned Geometric Latent Prediction

Jiangtao Liu*Zishang Xiang*Yage HeLingguo CuiBaihai ZhangRunqi ChaiSenchun Chai†

* Equal contribution.   † Corresponding author.

Method Overview

Figure 1. ACG-WAM combines geometric prediction conditioned on actions with distillation across camera views, evaluated in simulation and on a real robot.

Geometric Prediction Conditioned on Actions

ACG-WAM learns to predict future geometric representations from current observations and intervening actions. Our ACG-JEPA objective combines prediction over multiple horizons with geometric distillation from head and wrist cameras to train the policy’s shared visual embedding.

We evaluate ACG-WAM on 50 RoboTwin 2.0 tasks and three real robot tasks. Ablations on six simulation tasks examine the joint geometric targets, action conditioning, and supervision at multiple horizons.

Figure 2. ACG-JEPA supervises the MoT backbone’s shared visual embedding by predicting geometric targets from head and wrist cameras, conditioned on actions over multiple horizons.

ACG-WAM Architecture

Prediction conditioned on actions. ACG-JEPA predicts geometric targets from features of the current frame, intervening actions, and a temporal horizon. A frozen VGGT teacher jointly encodes pairs of current and future images and supplies targets from the future slot at multiple horizons.

Geometric distillation across views. Targets from head and wrist cameras supervise the MoT backbone’s shared visual embedding before temporal mixing, so the predictor’s visual input contains only the current observation. The geometric loss updates this embedding alongside the video and action objectives; the teacher and auxiliary predictor are removed at inference.

Real Robot Demonstrations

ACG-WAM performs bimanual fruit placement, block stacking, and toy placement into a cup on TRON2 with WUJI hands. Block stacking and toy placement are shown with both the robot’s left and right arms.

Bimanual Fruit Placement

Block Stacking

Toy Placement into a Cup

RoboTwin Simulation

Demonstrations across six manipulation tasks in clean and randomized scenes.

Clean scenes, demonstration 1

Citation

@unpublished{liu2026acgwam,
  title  = {ACG-WAM: World Action Modeling via Action Conditioned Geometric Latent Prediction},
  author = {Liu, Jiangtao and Xiang, Zishang and He, Yage and Cui, Lingguo and Zhang, Baihai and Chai, Runqi and Chai, Senchun},
  year   = {2026},
  note   = {Manuscript},
  url    = {https://RoboOpus.github.io/ACG-WAM/}
}