Back to Blogs

Active Vision & Occlusion / Original EGO R9 field guide informed by EgoAVFlow / Sep 7, 2026

Keep the Task in View: Active-Vision Demonstrations with EGO R9

A manipulation demonstration loses value when the hand, object, or goal disappears behind a drawer, tool, or robot arm. Inspired by EgoAVFlow, this guide shows how R9 teams can collect visibility-aware first-person data and test it before scale-up.

Researcher wearing EGO R9 while opening a drawer and reaching for a partly occluded object beside an active-vision robot test station
TINTELE GLOBAL CO., LIMITED original AI-generated application illustration based on authentic EGO R9 product imagery

Opening a drawer hides its contents behind the front panel. Reaching into a bin places the wrist between the camera and the target. A tool, forearm, cabinet door, or robot link can cover the exact contact point that determines whether a grasp or insertion succeeds. For first-person data collection, visibility is therefore part of the action rather than a cosmetic property of the video.

The EgoAVFlow paper examines this problem by learning manipulation and active viewpoint control from human egocentric videos through a shared 3D-flow representation. Its system predicts robot actions, future scene motion, and camera trajectories, then refines the view using a visibility-aware objective. The authors report stronger performance than prior human-demonstration baselines under changing viewpoints. Those methods and results belong to EgoAVFlow; EGO R9 is a separate capture platform that can help teams build and evaluate visibility-rich human demonstrations.

A useful R9 protocol begins with tasks in which occlusion changes over time. Examples include opening a drawer and selecting an object, retrieving parts from a deep bin, pouring from an opaque container, inserting a connector behind a cable bundle, placing an item on a crowded shelf, or fastening a component beneath a cover. Each sequence should include moments when the target is fully visible, partly hidden, and revealed again through natural head or body movement.

R9's 120-degree wide-angle lens helps keep both hands, the active object, and surrounding geometry in the same first-person frame. That additional context can show why the wearer leaned, rotated, or moved closer. Before collection, teams should set the adjustable camera angle for the real working distance and verify the lowest reach, nearest contact, deepest drawer position, and most obstructed phase of the task.

Motion quality matters when the participant changes viewpoint quickly. R9 records 1080P global-shutter video at 30 FPS, with 60 FPS and 1920 by 1200 available as optional configurations. Global shutter helps prevent rolling-shutter geometry from bending drawer edges, tools, or object contours during head turns. Exposure blur can still hide fine contact, so the pilot should test rapid look-around motions and hand transfers under the dimmest approved lighting.

Camera-intrinsics support and a consistent 120-degree view help teams preserve geometric context across repeated visibility trials. Record a calibration target at the intended working distance, store the intrinsics reference with the session, and inspect whether task-critical regions remain visible through each viewpoint change.

Head movement is also a signal. R9's 6-axis IMU samples above 200 Hz, recording acceleration and rotation between video frames. Shared clock support and global timestamps help align this motion with image sequences and approved task markers. Reviewers can compare a deliberate lean or turn with the moment an occluded object reappears, while still validating IMU continuity, axis conventions, timestamp monotonicity, and synchronization behavior on the ordered hardware.

A visibility-aware capture script should label more than success or failure. Mark when the target first becomes visible, when contact begins, when the view is blocked, when the participant changes viewpoint, when visibility is recovered, and whether the action then succeeds. This creates evidence for studying the relationship between perception and manipulation. H.265 recording in an MP4 container and optional approved microphone markers can support session handling, but annotation definitions must be designed by the dataset team.

Acceptance testing can make occlusion measurable. Define one or more task-critical regions, then calculate or manually score how often they remain visible during approach, contact, transport, and completion. Repeat the task with different participants, drawer depths, bin walls, object sizes, shelf heights, lighting, dominant hands, and safe recovery strategies.

Field operation must protect the same quality standard. Type-C supports live setup and bench inspection, while T-Flash storage supports untethered recording. The replaceable strap, adjustable angle, and external-battery support suit longer studies; the product specification states approximately five working hours with an external battery. Runtime, storage capacity, thermal behavior, comfort, strap stability, and recovery after interruption should all be verified under the chosen recording configuration.

The protocol should also define consent, approved areas, excluded screens and documents, bystander handling, retention, access, and stopping rules.

EGO R9 preserves first-person global-shutter imagery, wide scene context, high-rate inertial motion, and a shared time basis from human demonstrations.

The practical takeaway is simple: design demonstrations that challenge visibility on purpose. Record drawers, bins, shelves, covers, two-handed actions, and changing camera poses; validate image geometry and timing; annotate loss and recovery of task-critical views; and review the data with the intended downstream tools. With a disciplined protocol, R9 can help teams capture the relationship among seeing, moving, and manipulating instead of delivering a collection of attractive but incomplete POV clips.

active visionocclusion3D motionEGO R9
Explore more EGO field guidesDiscuss a camera data project