We present Act2See, a closed-loop system for reconstructing an explicit URDF of an unknown articulated object during interaction. A VLM planner uses the current RGB-D observation together with the maintained URDF state, joint list, and attempt history to decide whether to probe, execute, or skip actions. The system then grounds actions into grasps and constrained motions, estimates joint structure mainly from end-effector trajectories, and updates the URDF online. The main idea is that the URDF acts both as the reconstruction output and as structured memory for the planner.
This project website is under construction, the full project and arxiv will be released soon.