Research · Ember Lab, UC Berkeley · Fall 2026–Present
Teaching a ping-pong robot to see for itself
In the Ember Lab at Berkeley (affiliated with BAIR), I’m working on onboard vision tracking for the humanoid table-tennis robot from HITTER and LATTE-MV, so it can track the ball without relying on a motion-capture room.
The robot
The robot is a Unitree G1 humanoid, the same setup behind HITTER, which rallied 106 consecutive shots against a human player. HITTER splits the problem in two: a model-based planner predicts the ball’s flight, including drag and table bounces, and decides where, when, and how fast to strike; a reinforcement-learning controller trained with PPO and human forehand and backhand swing references moves the arms and legs to make the hit. LATTE-MV, from the same group, reconstructs table-tennis matches in 3D from ordinary monocular video and uses them to teach a policy to anticipate an opponent’s shot.
The problem
All of that depends on knowing where the ball is. In HITTER, that comes from nine OptiTrack motion-capture cameras running at 360 Hz, tracking a ball wrapped in reflective tape to millimeter accuracy. The paper names this dependence on external motion capture as a limitation: the robot can only play inside a room built around it.
What I’m working on
Giving the robot its own eyes. The goal is to estimate the ball’s position and velocity from onboard vision, quickly and accurately enough that the planner can still predict the trajectory and choose a hit, so the robot no longer needs motion capture to decide what to do. A table-tennis ball is small, fast, and blurry on camera, and a camera mounted on a moving humanoid moves with every step and swing, so there is very little room for latency or error.