Learning Reusable Hybrid Motion Priors for Humanoid Locomotion from Motion Imitation

Valerio Belli1,2 Valerio Modugno2 Enrico Mingo Hoffman3 Fabio Amadio3

1Sapienza University of Rome 2University College London 3Inria, Université de Lorraine, CNRS

We distill a motion-imitation expert into a frozen continuous-discrete Hybrid Motion Prior (HMP). Task-level policies control the robot by selecting entries from its codebook.

Watch video Code coming soon
Three-stage pipeline: motion imitation, Hybrid Motion Prior distillation, and task learning by code selection

Phase I

Motion imitation

A PPO expert tracks retargeted LaFAN motions, including walking, running, jumping, and recovery clips.

Phase II

HMP distillation

The expert is distilled into a proprioceptive encoder, an RVQ codebook, and an action decoder. These components form the frozen HMP.

Phase III

Task learning

For each downstream task, a task-level policy selects RVQ codebook entries while the HMP parameters remain frozen.

Evaluation

Downstream locomotion tasks

We train three task-level policies using the same frozen HMP.

01

Velocity tracking

The policy tracks planar and yaw-velocity commands and generates standing, walking, and running motions.

02

Point-goal navigation

The policy receives a target position relative to the robot and walks through consecutive goals.

03

Fall-recovery velocity tracking

The policy is trained from fallen configurations and under strong external perturbations, then resumes velocity tracking after recovery.

Video

Supplementary video

The video shows expert imitation, HMP distillation, the three downstream tasks, HMP analysis, and deployment on the real Unitree G1.