Back to Blogs

Research / arXiv / Feb 22, 2026

Human-to-Robot Interaction: Learning from Video Demonstration for Robot Imitation

This Human-to-Robot pipeline separates video understanding from robot imitation, turning observed actions and objects into structured manipulation skills.

Video understanding converted into robot manipulation skill diagram
TINTELE GLOBAL CO., LIMITED editorial context image

The Human-to-Robot study proposes a modular route from unstructured video demonstrations to robot manipulation. The workflow first interprets the demonstrated action and interacted objects, then passes that structured result to a robot-learning stage.

Its video-understanding module combines Temporal Shift Modules with vision-language models. The robot-imitation stage uses TD3-based deep reinforcement learning to execute reach, pick, move, and place actions.

The paper evaluates the method in PyBullet with a UR5e manipulator and in a real-world experiment with a UF850 manipulator. It reports 89.97 percent action-classification accuracy and an average robot-manipulation success rate of 87.5 percent across the evaluated actions.

The modular design makes task boundaries, object identity, action order, and visible state changes important properties of the source video. Those elements give the understanding stage clearer evidence to pass into robot execution.

The cited arXiv record contains the complete model design, baseline comparisons, evaluation protocol, and project resources.

video demonstrationrobot imitationvision-language modelsreinforcement learning
Explore more EGO field guidesDiscuss a camera data project