Back to Blogs

Research / arXiv / Jul 16, 2026

AeroAct: Action-Centered World-Action Models for Language-Conditioned Quadrotor Flight

AeroAct predicts smooth quadrotor trajectory-action chunks from egocentric visual history, proprioception, and language-conditioned goals.

AeroAct: Action-Centered World-Action Models for Language-Conditioned Quadrotor Flight related egocentric vision research visual
TINTELE GLOBAL CO., LIMITED editorial context image

AeroAct addresses language-conditioned quadrotor flight under rapidly changing first-person views. The model must connect a semantic goal with visual history and produce smooth control references that a physical aircraft can execute.

Its action-centered world-action model adapts a pretrained video diffusion Transformer to predict local trajectory-action chunks from egocentric visual history, proprioception, and language.

During training, future first-person frames provide dense supervision about the visual consequences of motion. During deployment, the system directly decodes actions from the learned representation.

The authors also introduce a handheld collection device that couples camera observations with motion estimates to reproduce flight-like egocentric trajectories. Simulation and real-world experiments report benefits from temporal visual context in target tracking and object search.

The cited arXiv record provides the complete model, data-generation pipeline, flight experiments, and reported results.

egocentric visionquadrotor navigationworld-action modelslanguage-conditioned control
Explore more EGO field guidesDiscuss a camera data project