EgoLAP

Learning from Egocentric Human Data
through Language-Action Reasoning

1 Princeton University2 Toyota Research Institute3 Physical Intelligence

* Equal contribution   ·   † Equal advising

01   /   Overview and method

Different bodies
A shared language of motion

EgoLAP learns from egocentric human and robot trajectories by expressing motion intent in language. Structured language actions describe what to do; motion-level reasoning explains why the motion fits the scene.

TrainingHuman + robot trajectories · ~10,000 hours
Training batch mixtureEgoVerse human data 25 percent, ABC YAM 30 percent, other bimanual robots 25 percent, OXE single-arm robots 20 percent.25%30%25%20%~10k hoursHUMAN + ROBOT DATA

Share of each training batch

EgoVerse

25%
Representative MECKA training camera viewMECKA 20%
Representative Scale training camera viewScale 4.15%
Representative Aria training camera viewAria 0.85%

ABC YAM

30%
Representative ABC YAM training camera viewABC YAM 30%

Other bimanual

25%
Representative AgiBot training camera viewAgiBot 15%
Representative Galaxea training camera viewGalaxea 5%
Representative MolmoAct2 training camera viewMolmoAct2 5%

OXE

20%
Representative DROID training camera viewDROID 12%
Representative Fractal training camera viewFractal 5%
Representative Bridge training camera viewBridge 2%

Other OXE datasets: 1%

Different motions can serve the same intent

Humans and robots can use different motions to achieve the same physical effect. The underlying intent is what generalizes. To generalize across embodiments, models need to understand why a motion works. Direct behavior cloning can leave this signal too implicit for the model to extract. Motion-level reasoning makes it explicit by connecting contact, geometry and object affordances to the intended effect.

Model Architecture

Training only

Motion-level reasoning

“Maintain the grip to carry one side over the other.”

Training only

Language actions

“Move forward 4 cm, move down 1 cm…”

Training + inference

Continuous actions

End-effector poses + gripper commands

Vision–language backboneLearn motion semantics across embodiments
Action expertFlow matching
Shared observation prefix
Prefix attention
Example human cloth-folding observationExample robot cloth-folding observation
Images · instruction · state“Fold the cloth into a square”

Example input from either embodiment

Illustrative example

Fold the cloth into a square

Different trajectories; shared intent: secure the cloth and bring one half over the other.

HumanEgocentric view
Human cloth folding: alignAlign
Human cloth folding: liftLift
Human cloth folding: foldFold
Human cloth folding: pressPress
Motion-level reasoning

The left hand maintains its grip and shifts rightward, bringing the left side over while the right hand anchors the opposite fold.

Language action
L: move forward 4 cm, move down 1 cm…
R: move back 5 cm, move down 4 cm…
RobotThird-person view
Robot cloth folding: alignAlign
Robot cloth folding: liftLift
Robot cloth folding: foldFold
Robot cloth folding: pressPress
Motion-level reasoning

With both corners firmly grasped, the arms lift upward and advance to fold the lower half over the upper half.

Language action
L: close gripper, move forward 7 cm…
R: move forward 7 cm, move up 4 cm, keep gripper closed…

02   /   Evaluation

Zero-shot performance on
unseen robot configurations

We evaluate on a customized bimanual YAM robot with unseen arm placement and camera configuration, and on 14 simulated tasks. No model is trained on simulation data. EgoLAP reaches 80.1% mean task progress in the real world and 38.3% success in simulation.

80.1%

Mean real-world task progress

2.3×

Gain over alternative action representations

38.3%

Zero-shot simulation success

Original full-resolution photo of the bimanual YAM evaluation setup
Real world   Unseen arm placement and camera configuration
Original simulation scene with two robot arms and manipulation objects
Simulation   14 tasks, no simulation training data

Real-world task progress

Open original ↗
Original real-world progress chart from Keynote showing EgoLAP, ABC-VLA, EgoFAST, and OAT with reasoning ablations
Mean task progress across five real-world behaviors. Error bars are retained from the original chart.

Zero-shot simulation success

Open original ↗
Original simulation success chart from Keynote comparing overall, pick, place, color, and fold task groups
Simulation success rate is a different metric from real-world task progress.
Vector chart: simulation success improves from 7.2 percent with robot data only to 16.4 percent with human and robot data

Human data makes a difference

With the language-action representation held fixed and no reasoning supervision, adding egocentric human data raises simulation success from 7.2% to 16.4%—a 2.3× gain.

The simulator is unseen during training, making this a direct test of transfer across both embodiment and visual domain.

03   /   Real-world rollouts

Zero-shot evaluation videos

01 / 06Fold pants
0:00

Zero-shot on an unseen YAM configuration, with no fine-tuning on the evaluation setup. Select a task to explore the rollouts.

04   /   Beyond pre-training

Effective Post-training

Vector chart: relative GRPO improvements are 19.4 percent with reasoning pre-training and language-action sampling, 12.7 percent when also sampling reasoning, and 14.2 percent without reasoning pre-training

Reinforcement learning on language actions

Structured language actions are machine-parseable and can be scored against demonstrated motion. GRPO post-training improves simulation performance while keeping the vision encoder and continuous action expert frozen.

Sampling language actions without generating reasoning gives the stronger improvement in this experiment. Deployment continues to use continuous actions without decoding text.

Citation

@misc{zha2026egolaplearningegocentrichuman,
  title={EgoLAP: Learning from Egocentric Human Data through Language-Action Reasoning},
  author={Lihan Zha and Shresth Grover and Tenny Yin and Samuel M. Bateman and Hengkai Pan and Mengchao Zhang and Aykut Onol and Allen Z. Ren and Dhruv Shah and Anirudha Majumdar},
  year={2026},
  eprint={2610.08726},
  archivePrefix={arXiv},
  primaryClass={cs.RO},
  url={https://arxiv.org/abs/2610.08726},
}