Act2See: Structured Articulated Representations through Robot Interaction

Yifeng Liu1   Yuzhen Chen2   Bingyang Wang3   Chengqi Li4   Jingyi Lu3   Han Li3   Kaichen Zhou1,2   Mengyu Wang2   Fangneng Zhan3
1Massachusetts Institute of Technology (MIT)   2Harvard University   3The Hong Kong University of Science and Technology (HKUST)   4Eindhoven University of Technology
accepted by RSS 2026 FM4RoboPlan Workshop
ABSTRACT

We present Act2See, a closed-loop system for reconstructing an explicit URDF of an unknown articulated object during interaction. A VLM planner uses the current RGB-D observation together with the maintained URDF state, joint list, and attempt history to decide whether to probe, execute, or skip actions. The system then grounds actions into grasps and constrained motions, estimates joint structure mainly from end-effector trajectories, and updates the URDF online. The main idea is that the URDF acts both as the reconstruction output and as structured memory for the planner.

This project website is under construction, the full project and arxiv will be released soon.

Real-world Deployment Demonstration
20260706_005854_drawer_pull
20260706_005423_drawer_pull
20260705_210432_door_e2b
20260705_211145_door_e2b