Research / arXiv / Jul 17, 2026
Exo2EgoPose: Leveraging Exocentric Demonstrations for Vision-Language guided Egocentric 3D Hand Pose Forecasting
Exo2EgoPose uses exocentric demonstrations to guide vision-language prediction of future 3D hand poses from dynamic egocentric observations.
Exo2EgoPose introduces vision-language-guided egocentric 3D hand-pose forecasting. The task predicts future hand poses from first-person visual observations, a language instruction, and current pose states.
Dynamic motion and partial views make egocentric forecasting difficult. The framework uses stable, wider exocentric demonstrations as guidance for spatial context and temporal development.
A dual-level exocentric reconstruction module learns video-level and chunk-level representations. A global-to-local modulation module then uses those reconstructed features to refine the egocentric prediction.
The paper evaluates the method on AssemblyHands, Ego-Exo4D, and the authors' EgoMe-pose benchmark, reporting substantial improvements over the evaluated methods.
The cited arXiv record contains the full architecture, benchmark construction, experimental comparisons, and updated paper versions.
