Phase I
Motion imitation
A PPO expert tracks retargeted LaFAN motions, including walking, running, jumping, and recovery clips.
1Sapienza University of Rome 2University College London 3Inria, Université de Lorraine, CNRS
We distill a motion-imitation expert into a frozen continuous-discrete Hybrid Motion Prior (HMP). Task-level policies control the robot by selecting entries from its codebook.
Phase I
A PPO expert tracks retargeted LaFAN motions, including walking, running, jumping, and recovery clips.
Phase II
The expert is distilled into a proprioceptive encoder, an RVQ codebook, and an action decoder. These components form the frozen HMP.
Phase III
For each downstream task, a task-level policy selects RVQ codebook entries while the HMP parameters remain frozen.
Evaluation
We train three task-level policies using the same frozen HMP.
The policy tracks planar and yaw-velocity commands and generates standing, walking, and running motions.
The policy receives a target position relative to the robot and walks through consecutive goals.
The policy is trained from fallen configurations and under strong external perturbations, then resumes velocity tracking after recovery.
Video
The video shows expert imitation, HMP distillation, the three downstream tasks, HMP analysis, and deployment on the real Unitree G1.