akina
PLATE 02 · YUME

Train your VLA on human demonstrations.

Yume turns any RGB video of a person at work into action-labelled training data: wrist and finger trajectories in 3D, in meters, a language instruction for every action, and the same motion solved onto your robot. Pretrain on what people already do, then fine-tune on your robot.

Talk to usAvailable · runs in the cloud
left hand

Wipe the steam wand of the coffee machine with a cloth.

right hand

Wipe the steam wand with the cloth.

Task
The person is making a flat white coffee.
Summary
The person wipes the coffee machine with a cloth, then pours milk from a bottle into a jug. They place a saucer on the counter, take a cup of coffee from the machine, and pour the milk into the coffee.
Objects
clothcoffee machinemilk jugmilk bottlesaucercupcoffeemilk
WHY HUMAN VIDEO

Human motion scales like text.

Recent work¹ pretrained a VLA on 20,854 hours of action-labelled egocentric human video, then fine-tuned it on a robot. Validation loss fell log-linearly with hours of human video and tracked success on the real robot. The same pretraining enabled one-shot learning of new tasks and transferred to a Unitree G1 with a lower-DoF hand. What made it work was the labels: 3D wrist and hand actions for every frame. Yume produces those labels from any RGB video, first or third person.

20,854 h
of action-labelled human video in pretraining
log-linear
validation loss against hours, tracking robot success
+54%
average success on a 22-DoF hand over no human pretraining
MORE THAN HAND TRACKING

The whole scene, in 3D, in meters.

Yume does not stop at the hands. From the same video it recovers the camera's path, the surfaces and objects around the person and the wearer's whole body, all in one metric world frame. With the scene in 3D, one video feeds three more things.

Café · first person · scene, camera path and wearer in 3D
RETARGETING

The same motion, solved onto your robot.

Every episode also comes with the motion on your robot's joints: within joint limits, without self-collision and with feet on the ground. The same pipeline holds across settings, here an orchard, a café, an engine line and a building site, all solved onto a G1. G1, H1, H1-2, T1, A3 and dexterous hands today; a new robot is a config, not code.

HAND SWAP

Swap the human hand for a robot hand.

Yume's 3D hands, projected through the estimated camera, give a mask on every frame. Hide the arm, blur it, or inpaint it and render the robot's hand from the solved trajectory in its place, so human video reads like robot data to a policy. Available on request.

OriginalSegmentBlurInpaint + robot
DIGITAL TWINS

Filmed once, simulated many ways.

Yume's reconstruction becomes a CAD model of the workspace, within 3 cm where the work happens. Swap tools, parts and layout to make its cousins, then replay the task under physics and keep what succeeds: one recording, many scenes to train on. Available on request.

WHAT YOU GET

Outputs that fit the tools you already use.

Episodes
One action per episode: the language instruction, the objects it names, video, and wrist pose and finger joints over time in the world frame. LeRobot format, ready for VLA pretraining or fine-tuning.
Robot
Joint trajectories for G1, H1, H1-2, T1, A3 and dexterous hands, each with a quality report.
Human
Whole body and 21 joints per hand with a rigged mesh, in meters.
Scene
Camera intrinsics and path, gravity, metric depth and a static point cloud, in one world frame.
HOW TO START
  1. Send a video
    Any clip of the task: head camera, phone or fixed camera.
  2. Pick your robots
    Choose the embodiments. Each gets its own solve and quality report.
  3. Review and train
    Inspect episodes in the 3D viewer, then download the dataset and train your VLA on it.
QUESTIONS
Which VLAs?
Any model that trains from a LeRobot dataset. Each episode carries its video, actions and language instruction.
Do I still need robot demos?
A few. Human video teaches the motion and a small set of demos on your robot grounds it in your hardware. In published results¹, human pretraining raised success for the same robot fine-tuning data and made one-shot learning of new tasks possible.
What video works?
First or third person, from a head camera, phone or fixed camera. Hands in frame for manipulation, whole body in frame for body motion.
Which robots?
G1, H1, H1-2, T1, A3 and dexterous hands today. A new robot is a config, not code: send us the URDF. Every solve comes with a quality report.

If it was filmed, a robot can learn it.