As robots increasingly operate in human-populated environments, anticipating human intentions is essential for enabling proactive and socially aware behavior. Automatic anticipation of human–robot interactions is thus emerging as a crucial perception challenge for embodied agents.
To this end, we introduce HUI360, the largest dataset for human-robot interaction anticipation in the wild and its set of baselines. The dataset was collected from a mobile robot, in the wild, over multiple days within a 3-month period, and in several environments, capturing natural, spontaneous behaviors from both passersby and users, and encompassing a diverse range of individuals. This variety enables evaluating and improving the generalization capabilities of interaction anticipation models.
We designed a pipeline and share code for automatic interaction annotation in arbitrary 360° equirectangular videos, along with interfaces for manual refinement. Using this pipeline, we release the HUI360 open set of 1M pre-processed annotations, including detailed 2D poses, facial keypoints, and segmentation masks, obtained using state-of-the-art computer vision methods and manually curated to ensure high-quality tracking and interaction annotation. Additionally, we release the raw panoptic 360° images captured from the robot’s egocentric viewpoint (on demand, for research purpose only in compliance with GDPR).
Finally, we establish benchmark baselines for interaction anticipation, including the first cross-dataset evaluations for this task: to this end, we also release 6M annotations for another existing in-the-wild outdoor dataset collected from a mobile robot (SSUP-HRI).
(Recordings split) |
HUI360 (Ours - Indoor w/ Shelfy) |
HUI360 (SSUP-HRI - Outdoor w/ Trashcan) |
|---|---|---|
| Recording duration | 71h | 26h |
| Recording duration (after filtering) | 11h | — |
| Individual tracks | 4310 | 28000 |
| Interactions with the robot | 375 | 419 |
| Images | 621,000 | 1.4M |
| Detections | > 1M | 6M |
Modalities comparison (better seen with Chrome).
You can freely access the processed annotations (skeletons, keypoints, masks, etc.) on HUI360.
Access to the raw 360° videos from our Shelfy recordings requires a Data Transfer Agreement (DTA). These DTAs cover our HUI360 recordings only — not the SSUP-HRI videos, which are subject to their own agreements (see the SSUP-HRI section). Access is reserved for researchers at academic institutions, for research purposes only, and subject to GDPR compliance.
To request access to the raw videos:
Note: As of October 16, 2023, institutes located in the following territories may use the EEA / adequate-countries form: countries of the European Union, Iceland, Liechtenstein, Norway, Andorra, Argentina, Canada, Faroe Islands, Guernsey, Israel, Isle of Man, Japan, Jersey, New Zealand, Switzerland, Uruguay, South Korea, and the United Kingdom. All other affiliations should use the outside-EEA form.
We also provide the pipeline to automatically annotate 360° videos of human-robot interaction from the robot's egocentric viewpoint. You can run it with videos from cameras like the Insta360 attached to a robot or a person.
We recorded 71h of multimodal data and kept 11h of recordings with passerbys. Main data available are processed from 360° equirectangular videos.
If you are interested in other modalities please contact us (not available for all sessions).
Sensorized Shelfy robot.
Part of the HUI360 dataset is based on the amazing SSUP-HRI dataset from Cornell IRL Team. To access the SSUP-HRI dataset please visit SSUP-HRI and request access following the instructions on the repository.
@inproceedings{lorenzolouis:hal-05609928,
TITLE = {{HUI360 : A 360{\textdegree} Egocentric Dataset and Baselines for Human-Robot Interaction Anticipation}},
AUTHOR = {Lorenzo-Louis, Raphael and Amadio, Fabio and Luvison, Bertrand and Ivaldi, Serena},
URL = {https://hal.science/hal-05609928},
BOOKTITLE = {{FG 2026 - IEEE International Conference on Automatic Face and Gesture Recognition}},
ADDRESS = {Kyoto, Japan},
ORGANIZATION = {{IEEE Biometrics}},
YEAR = {2026},
MONTH = May,
KEYWORDS = {Human robot Interaction ; Computer vision ; Robotics ; AI ; Vision},
PDF = {https://hal.science/hal-05609928v1/file/FG_2026_AuthorVersion_InclSupp_compressed.pdf},
HAL_ID = {hal-05609928},
HAL_VERSION = {v1},
}