OCC4M
Supplementary videos

Object-Centric 4D Memory for Spatiotemporal Reasoning in Long-Horizon Manipulation

Remembering where objects were, which objects they are, and what they contain—then grounding that history for manipulation.

Simulation experiments

Each video illustrates a task requiring information from earlier observations. Watch the initial history before the robot acts; use slower playback to inspect brief events.

T1

Spatial persistence

Open video

Task. Pick the red cube and place it where the green cube was.

What to watch. Watch the placement target remain tied to the green cube’s earlier location after it disappears.

T1c

Viewpoint transfer

Open video

Task. Pick the red cube and place it where the green cube was, after the green cube disappears and the robot base moves.

What to watch. Watch the remembered location be grounded from the robot’s new viewpoint.

T2

Temporal identity

Open video

Task. Place the blue cube at the location of the marker that appeared first, and the red cube at the location of the marker that appeared second.

What to watch. Watch the order in which the markers appear; that order determines the two placement targets.

T3

Event history

Open video

Task. Pick the second cube that moved and place it at the first mover’s pre-motion location.

What to watch. Watch which cube moves first and where it was before moving.

T4

Relational persistence

Open video

Task. Pick the grey cover containing the red cube after the covers are shuffled.

What to watch. Follow the cover–object relation through the shuffle, while the cube is hidden.

T5

Cross-view association

Open video

Task. Pick the green cube from the second table and place it where the green sign appeared on the first table.

What to watch. Watch the target location be recalled across views of the two tables. The query color is randomized per episode; this video shows green.

T6

Multi-instance identity

Open video

Task. Drop every cube that was originally red into the left bin and every cube that was originally green into the right bin.

What to watch. Watch the original identities determine sorting after the cubes change appearance.

Real-world experiments

Franka demonstrations of compositional recall and online target grounding.

Franka

Compositional manipulation

Open video

Task. Pick two cups in this order: (1) the cup covering the red cube, and (2) the cup occupying the white mug’s original location.

What to watch. The two-stage Franka task combines shuffled-containment recall with historical-location recall.

Franka

Closed-loop grounding under perturbation

Open video

Task. Pick the cup covering the red cube. During execution, the target cup is displaced.

What to watch. Watch the selected target follow the displaced cup without a new VLM query. This is a single perturbation demonstration.