Research / arXiv / Jul 16, 2026
Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation
Ego Scene Augmentation builds an Ego-element Graph to strengthen spatial reasoning by multimodal language models in first-person scenes.
Ego Scene Augmentation focuses on spatial reasoning in egocentric visual question answering. First-person scenes can contain close objects, partial views, changing orientation, and relationships that are difficult to express from pixels alone.
The framework constructs an Ego-element Graph that integrates spatial features supplied by visual foundation models. This graph becomes an intermediate representation for strengthening an MLLM's understanding of ego-perspective scenes.
The method uses the graph to connect scene elements and guide spatial reasoning. It is designed to retain the relationships that matter when a first-person viewer asks where objects are and how they relate to the current viewpoint.
On EgoTextVQA, the paper reports gains of 8.14 percent in indoor settings and 8.72 percent outdoors, with its strongest improvement in the indoor shopping subset.
The cited arXiv publication contains the graph construction, augmentation method, benchmark protocol, ablations, and code link.
