SPOT combines an immersive perception interface with a sparse VR control pipeline. A robot-mounted binocular fisheye camera streams stereo observations onto a virtual hemisphere around the operator, preserving depth cues and peripheral coverage. In parallel, tracked headset and controller poses are used to estimate the torso and map wrist targets to the robot, while joystick inputs command locomotion.
Abstract
Long-horizon humanoid teleoperation requires sustained awareness of objects, the environment, and the robot’s body. SPOT extends this perceptual horizon with a wide-field binocular view rendered on a virtual hemisphere. Operators can look around without directly commanding robot motion, while visual stabilization maintains a consistent spatial reference. Across perception-critical tasks, SPOT improves task efficiency over conventional egocentric teleoperation.
Video
Method

Viewpoint-decoupled free-looking. Head rotations let the operator inspect different parts of the wide-field observation without directly turning the robot’s head or torso. A learned torso estimator separates body commands from natural head motion, allowing the operator to look around while continuing to manipulate.
Visual stabilization. An IMU on the camera rig tracks its orientation. The virtual display compensates for camera rotations caused by locomotion and balancing, helping the operator maintain a consistent spatial reference as the robot moves.

Demonstrations
Continuous locomotion and manipulation, from the operator’s view and an external view.
Interact with moving targets and coordinate both hands across a wide workspace.
Tie knots while keeping both hands and the rope in view.
Comparison
A wider view helps the operator track peripheral objects and both manipulators without repeatedly searching or switching focus.
SPOT
Conventional egocentric
Results
SPOT achieves the lowest mean task completion time across all four tasks, with success rates of 98–100%. Compared with conventional egocentric teleoperation, it reduces recovery time from 20.28 to 11.12 seconds and alignment time from 1.52 to 0.39 seconds, while requiring fewer grasp attempts during bimanual retrieval.

BibTeX
@misc{fang2026spotspatialperceptionorientedlonghorizon,
title={SPOT: Spatial Perception-Oriented Long-Horizon Humanoid Teleoperation},
author={Lixing Fang and Ziyan Xiong and Sunli Chen and Zhiyang Dou and Chuang Gan},
year={2026},
eprint={2609.07933},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2609.07933},
}