SPOT: Spatial Perception-Oriented Long-Horizon Humanoid Teleoperation

*Equal contribution
Conference on Robot Learning (CoRL), 2026

See more. Look freely. Stay oriented.

SPOT interface and demonstrations of knot tying, dynamic interception, recovery, and bimanual manipulation.

SPOT enables long-horizon humanoid teleoperation through wide-field stereo vision, viewpoint decoupling, and visual stabilization.

Abstract

Long-horizon humanoid teleoperation requires sustained awareness of objects, the environment, and the robot’s body. SPOT extends this perceptual horizon with a wide-field binocular view rendered on a virtual hemisphere. Operators can look around without directly commanding robot motion, while visual stabilization maintains a consistent spatial reference. Across perception-critical tasks, SPOT improves task efficiency over conventional egocentric teleoperation.

Video

Method

SPOT combines an immersive perception interface with a sparse VR control pipeline. A robot-mounted binocular fisheye camera streams stereo observations onto a virtual hemisphere around the operator, preserving depth cues and peripheral coverage. In parallel, tracked headset and controller poses are used to estimate the torso and map wrist targets to the robot, while joystick inputs command locomotion.

SPOT perception and control pipeline.

Viewpoint-decoupled free-looking. Head rotations let the operator inspect different parts of the wide-field observation without directly turning the robot’s head or torso. A learned torso estimator separates body commands from natural head motion, allowing the operator to look around while continuing to manipulate.

Visual stabilization. An IMU on the camera rig tracks its orientation. The virtual display compensates for camera rotations caused by locomotion and balancing, helping the operator maintain a consistent spatial reference as the robot moves.

Visual stabilization during operator and robot turns.

Demonstrations

Continuous locomotion and manipulation, from the operator’s view and an external view.

Interact with moving targets and coordinate both hands across a wide workspace.

Tie knots while keeping both hands and the rope in view.

Comparison

A wider view helps the operator track peripheral objects and both manipulators without repeatedly searching or switching focus.

SPOT

Conventional egocentric

Results

SPOT achieves the lowest mean task completion time across all four tasks, with success rates of 98–100%. Compared with conventional egocentric teleoperation, it reduces recovery time from 20.28 to 11.12 seconds and alignment time from 1.52 to 0.39 seconds, while requiring fewer grasp attempts during bimanual retrieval.

Table 1 comparing exocentric, egocentric, and SPOT teleoperation on unexpected drop, peripheral retrieval, bimanual retrieval, and light switch tasks. SPOT completion times are 19.69, 7.43, 9.76, and 11.54 seconds, with success rates of 98%, 100%, 99.5%, and 100%, respectively.

BibTeX

@misc{fang2026spotspatialperceptionorientedlonghorizon,
      title={SPOT: Spatial Perception-Oriented Long-Horizon Humanoid Teleoperation}, 
      author={Lixing Fang and Ziyan Xiong and Sunli Chen and Zhiyang Dou and Chuang Gan},
      year={2026},
      eprint={2609.07933},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2609.07933}, 
}