Skip to main content
GameDev.net gamedev.net
Research Paper

This is an academic paper or technical research. Key findings may require technical background to fully understand.

Explore Research Radar

PRO Tired of ads? Read GameDev.net ad-free and help keep the community independent with GameDev Pro — $3/month.

arXiv cs.GR
arXiv cs.GR Research
· 2 days, 3 hours ago • Haozhuo Zhang, Qiang Zhang, Jian Tang, Mingzhe Ni, Michele Caprio, Angelo Cangelosi, Wei Pan

Humanoid Horizon: Extending Task Horizon in Whole-Body Loco-Manipulation via Parallel Training, Dynamic Starting, and Reward Gating

Briefing

arXiv cs.GR details Humanoid Horizon, a new policy framework for humanoid robots that tackles long-horizon loco-manipulation in cluttered indoor spaces. The goal is to let a single agent navigate, grasp, carry, and place multiple heavy objects in sequence without earlier progress being undone as the episode continues.

The core idea is to split training across concurrent stage streams instead of optimizing each transport step in strict order. Parallel Training keeps gradients flowing to every stage, Dynamic Starting seeds environments from upstream terminal states to widen boundary coverage, and Reward Gating shuts off later-stage rewards when a previously placed object gets displaced.

That combination is meant to address two common failure modes in long-horizon robotics: easy-reward bias, where early stages dominate learning, and catastrophic forgetting, where later-stage optimization hurts earlier skills. For game developers, the interesting angle is less the robot itself than the training pattern: the paper is essentially about preserving multi-step state consistency under shared-policy learning.

The system reportedly reaches per-stage success rates above 80% on the two-object LHM-Humanoid benchmark, which uses 350 training scenes and 66 held-out scenes. Performance still drops as the number of sequential objects increases beyond two, but the decline is described as much smoother than the sharp collapse seen in the baselines.

“per-stage success rates exceeding 80%”

— Paper authors · Reported benchmark result
Original source
Read on arXiv cs.GR
At a glance
what
Humanoid Horizon is a unified policy framework for long-horizon humanoid loco-manipulation.
who
Paper by Haozhuo Zhang and six coauthors, published on arXiv cs.GR.
when
Submitted on 6 Oct 2026.
impact
Could inform robotics-style training loops for multi-step state preservation and shared-policy learning.
Signal Mixed

Promising results, but scaling still degrades with horizon

Discuss

Follow ai updates

See relevant stories in your personalized news feed.

Sign in to follow

Continue on GameDev.net

Useful next steps related to this story.

Game development news without the noise

One useful weekly briefing. No daily flood.

Sending your confirmation email…

Discussion

Loading comments...

Story Timeline (2 sources)

Story covered over 1 day • First reported by arXiv cs.GR