Humanoid Horizon: Extending Task Horizon in Whole-Body Loco-Manipulation via Parallel Training, Dynamic Starting, and Reward Gating
arXiv cs.GR details Humanoid Horizon, a new policy framework for humanoid robots that tackles long-horizon loco-manipulation in cluttered indoor spaces. The goal is to let a single agent navigate, grasp, carry, and place multiple heavy objects in sequence without earlier progress being undone as the episode continues.
The core idea is to split training across concurrent stage streams instead of optimizing each transport step in strict order. Parallel Training keeps gradients flowing to every stage, Dynamic Starting seeds environments from upstream terminal states to widen boundary coverage, and Reward Gating shuts off later-stage rewards when a previously placed object gets displaced.
That combination is meant to address two common failure modes in long-horizon robotics: easy-reward bias, where early stages dominate learning, and catastrophic forgetting, where later-stage optimization hurts earlier skills. For game developers, the interesting angle is less the robot itself than the training pattern: the paper is essentially about preserving multi-step state consistency under shared-policy learning.
The system reportedly reaches per-stage success rates above 80% on the two-object LHM-Humanoid benchmark, which uses 350 training scenes and 66 held-out scenes. Performance still drops as the number of sequential objects increases beyond two, but the decline is described as much smoother than the sharp collapse seen in the baselines.
“per-stage success rates exceeding 80%”
- what
- Humanoid Horizon is a unified policy framework for long-horizon humanoid loco-manipulation.
- who
- Paper by Haozhuo Zhang and six coauthors, published on arXiv cs.GR.
- when
- Submitted on 6 Oct 2026.
- impact
- Could inform robotics-style training loops for multi-step state preservation and shared-policy learning.
Promising results, but scaling still degrades with horizon
Follow ai updates
See relevant stories in your personalized news feed.
Discussion