Train your VLA on human demonstrations.
Yume turns any RGB video of a person at work into action-labelled training data: wrist and finger trajectories in 3D, in meters, a language instruction for every action, and the same motion solved onto your robot. Pretrain on what people already do, then fine-tune on your robot.
Wipe the steam wand of the coffee machine with a cloth.
Wipe the steam wand with the cloth.
- Task
- The person is making a flat white coffee.
- Summary
- The person wipes the coffee machine with a cloth, then pours milk from a bottle into a jug. They place a saucer on the counter, take a cup of coffee from the machine, and pour the milk into the coffee.
- Objects
- clothcoffee machinemilk jugmilk bottlesaucercupcoffeemilk
Human video, labelled as actions
Every clip returns wrist poses in a metric world frame, finger joints and a caption per action: the labels a VLA trains on.
Any camera, any setting
First or third person, head camera, phone or a fixed camera on the line. No capture rig, no markers, no calibration.
Your robots keep working
The dataset grows from footage, not teleoperation. Robot time goes to fine-tuning and evaluation.
Human motion scales like text.
Recent work¹ pretrained a VLA on 20,854 hours of action-labelled egocentric human video, then fine-tuned it on a robot. Validation loss fell log-linearly with hours of human video and tracked success on the real robot. The same pretraining enabled one-shot learning of new tasks and transferred to a Unitree G1 with a lower-DoF hand. What made it work was the labels: 3D wrist and hand actions for every frame. Yume produces those labels from any RGB video, first or third person.
- 20,854 h
- of action-labelled human video in pretraining
- log-linear
- validation loss against hours, tracking robot success
- +54%
- average success on a 22-DoF hand over no human pretraining
The whole scene, in 3D, in meters.
Yume does not stop at the hands. From the same video it recovers the camera's path, the surfaces and objects around the person and the wearer's whole body, all in one metric world frame. With the scene in 3D, one video feeds three more things.
The same motion, solved onto your robot.
Every episode also comes with the motion on your robot's joints: within joint limits, without self-collision and with feet on the ground. The same pipeline holds across settings, here an orchard, a café, an engine line and a building site, all solved onto a G1. G1, H1, H1-2, T1, A3 and dexterous hands today; a new robot is a config, not code.
Swap the human hand for a robot hand.
Yume's 3D hands, projected through the estimated camera, give a mask on every frame. Hide the arm, blur it, or inpaint it and render the robot's hand from the solved trajectory in its place, so human video reads like robot data to a policy. Available on request.
Filmed once, simulated many ways.
Yume's reconstruction becomes a CAD model of the workspace, within 3 cm where the work happens. Swap tools, parts and layout to make its cousins, then replay the task under physics and keep what succeeds: one recording, many scenes to train on. Available on request.
Outputs that fit the tools you already use.
- Episodes
- One action per episode: the language instruction, the objects it names, video, and wrist pose and finger joints over time in the world frame. LeRobot format, ready for VLA pretraining or fine-tuning.
- Robot
- Joint trajectories for G1, H1, H1-2, T1, A3 and dexterous hands, each with a quality report.
- Human
- Whole body and 21 joints per hand with a rigged mesh, in meters.
- Scene
- Camera intrinsics and path, gravity, metric depth and a static point cloud, in one world frame.
- Send a videoAny clip of the task: head camera, phone or fixed camera.
- Pick your robotsChoose the embodiments. Each gets its own solve and quality report.
- Review and trainInspect episodes in the 3D viewer, then download the dataset and train your VLA on it.
- Which VLAs?
- Any model that trains from a LeRobot dataset. Each episode carries its video, actions and language instruction.
- Do I still need robot demos?
- A few. Human video teaches the motion and a small set of demos on your robot grounds it in your hardware. In published results¹, human pretraining raised success for the same robot fine-tuning data and made one-shot learning of new tasks possible.
- What video works?
- First or third person, from a head camera, phone or fixed camera. Hands in frame for manipulation, whole body in frame for body motion.
- Which robots?
- G1, H1, H1-2, T1, A3 and dexterous hands today. A new robot is a config, not code: send us the URDF. Every solve comes with a quality report.