Human-Robot Interaction / Simulation

A humanoid that imitates the arms of a person seen by a camera

Jan 15, 20263 min read
Humanoid model in the MuJoCo simulation environment
Undergraduate thesis: Suárez Martínez, Jerónimo. Vision-based human motion imitation for humanoid robots. Universidad de los Andes, 2026.

A person moves their arms in front of a camera and a humanoid robot repeats the motion in real time, with no capture suit or markers. This project built that full chain for the Unitree G1 robot and validated it first in simulation and then on the physical robot. Across four test scenarios, mean tracking error per arm stayed below 3°.

Context

Teaching a humanoid a movement by demonstration is more natural than programming it joint by joint. The difficulty is that the robot's body is not the person's: its segments have other lengths and its joints other ranges. Copying angles directly does not work. The intent of the movement has to be transferred to the robot's kinematics, which is known as retargeting.

Architecture

System architecture

Data flow, from the human operator to the robot

Pose estimation. An RGB-D camera runs a body pose estimation model that delivers body landmarks in three dimensions.

Estimated skeleton

Upper-body skeleton estimated by the camera

Transformation and filtering. The landmarks are taken from the camera frame to a frame centered on the user and aligned with the robot's. A filter discards lost or unstable points.

Retargeting. The person's shoulder-elbow and elbow-wrist vectors are scaled to the robot's arm lengths. The resulting target position is converted into joint angles by inverse kinematics on the robot model in MuJoCo.

Arm vectors

Vector representation of the right arm after geometric scaling

Execution. A ROS 2 node validates the angles and sends them to the robot over Ethernet. Simulation and the real robot share the same command interface, so what is tested in simulation runs unchanged on the hardware.

Poses in simulation

Robot poses obtained in simulation from target points

Evaluation

Four scenarios were tested: T-pose, arms forward, an asymmetric greeting sequence, and free motion. Error was computed as the difference between commanded and measured angle at each joint.

Error in the greeting sequence

Mean absolute tracking error in the greeting sequence, per arm

Error summary

Mean tracking error per arm in each scenario

Scenario Right arm Left arm
T-pose 2.3° 2.1°
Arms forward 2.7° 2.4°
Greeting 2.6° 1.2°
Free motion 1.8° 1.7°

Error peaks appear at the start of each movement and reach about 6°. The robot's motion is smooth and visually consistent with the person's.

What is missing

The reported error measures how well the robot follows the commanded angles. It does not measure how similar the robot's pose is to the person's, which is the underlying question in imitation and would require an external reference. The user has to stay facing the camera, and estimation degrades with occlusions and fast movements. Only shoulders and elbows are imitated, with no wrist orientation or hands, which prevents manipulating objects. Learning action sequences remained early-stage work.

How it fits in Robiolab

This project and the one on reasoning with language models are the lab's first two works on the Unitree G1 humanoid, and they are complementary: one addresses how the robot decides what to do, and this one, how it acquires a movement from a person. Human motion capture with depth cameras is also a tool shared with the hand biomechanics projects, and imitation is a direct path toward teleoperation.

Humanoid robotsMotion imitationPose estimationInverse kinematicsROS 2