Everyday Life Data / Rice RobotPI Lab / EgoInfinity / Aug 17, 2026
From Meal Prep to 4D Robot Learning: Finding Structure in Everyday First-Person Video
Cooking, cleaning, organizing, and simple repairs contain rich contact and motion patterns. EgoInfinity illustrates how ordinary video can be lifted into structured 4D hand-object interaction data.
Preparing a simple meal involves more intelligence than it appears. A person locates ingredients, steadies a cutting board, changes grip for different objects, moves a utensil around obstacles, checks the state of the food, clears waste, and reorganizes the workspace. These small adjustments form a long sequence of perception, contact, motion, and verification.
EgoInfinity, introduced by Rice RobotPI Lab on June 27, 2026, is a web-scale engine for turning ordinary RGB video into structured 4D hand-object interaction data. Its pipeline seeks to recover hand and object geometry, contact, and motion over time, then retarget recovered human motion into robot joint trajectories. The project is built on Action100M, a large filtered collection of human-action clips.
The everyday-life connection is direct. Meal preparation includes cutting, pouring, stirring, opening, wiping, and placing. Household organization includes folding, stacking, sorting, loading, and carrying. Simple maintenance adds holding, aligning, tightening, testing, and putting tools away. Each activity contains relationships among hands, objects, space, and time that a flat action label cannot describe.
EgoInfinity emphasizes that internet-scale video offers diversity, but uncontrolled footage also varies in viewpoint, resolution, occlusion, and task completeness. A purpose-built first-person collection can complement web data with repeatable framing and known capture settings. It can also intentionally record the beginning, middle, completion, and safe recovery steps that short online clips often omit.
EGO R9 fits ordinary household and service-task collection as a hands-free first-person recorder. Its 120-degree wide-angle view helps retain both hands, the active object, and nearby context during cooking, cleaning, organizing, or tool use. Global-shutter 1080P video supports quick hand and head movement, while the 6-axis IMU above 200 Hz, shared clock, and global timestamps can help connect visual motion with task annotations or approved external sensors.
For example, a kitchen dataset could capture the complete process of washing vegetables, selecting a utensil, cutting, transferring ingredients, wiping the surface, and returning tools. A home-organization dataset might follow sorting laundry, folding items, opening storage, placing objects, and correcting a poor fit. The useful unit is the full task sequence, including pauses and ordinary adjustments, rather than only a successful final action.
The camera does not itself reconstruct 4D geometry, infer contact forces, or produce executable robot commands. Those are downstream research and processing steps. R9 provides the source visual and inertial evidence. Teams still need calibration, quality review, segmentation, privacy filtering, hand-object annotation, reconstruction, and validation against their intended robot embodiment.
Home and daily-life data requires strong consent and privacy controls. Recording plans should exclude private documents, screens, bystanders, addresses, and sensitive moments; specify retention and access; and use only safe, voluntary tasks. When those foundations are in place, ordinary first-person activity can become a valuable bridge between how people naturally handle the world and how future robots learn to assist with it.
