Robotics · Perception · Embodied AI

Zeyuan Gao

I build robot learning and perception systems for real-world deployment.

MRes Medical Robotics, Imperial College London
Embodied AI / Robot Learning Intern, Latencore AI

Real-robot deploymentDual-arm towel manipulation

Selected work · Robot learning

From demonstrations to dual-arm deployment

Deploying and debugging a vision-language-action policy on two AgileX PiPER arms.

Rigid-object pick-and-placeπ0.5 policy · dual-arm execution
Deformable towel manipulationπ0.5 policy · dual-arm execution
SYSTEM

More than a policy checkpoint

Real deployment connected teleoperation, multi-view demonstrations, data preparation, inference, communication and physical execution.

CONTRIBUTION

End-to-end integration

I worked across demonstration capture, end-effector trajectory preparation, LeRobot-compatible data, policy deployment and system-level debugging.

ENGINEERING

Failure diagnosis on hardware

I traced calibration, camera-stream, robot-communication, inference and action-handoff issues on the physical platform.

Data collection

Synchronized, multi-view demonstrations

PIKA capture combined two wrist cameras with an external panoramic view. I also integrated and operated a colleague-provided 3D-printed UMI setup using VIVE trackers and four Lighthouse base stations.

Contribution boundary. My work focused on integration, data workflows, deployment and debugging; the underlying robot controller and UMI hardware were developed by others.

PIKA three-view captureLeft wrist · panoramic view · right wrist

Selected work · Individual research

Geometry-aware reranking for zero-shot 6-DoF pose estimation

Imperial College London · Hamlyn Centre
Supervised by Dr Stamatia (Matina) Giannarou

RESEARCH PROBLEM

Foundation-model candidates can appear plausible while the default top-ranked pose remains geometrically inconsistent for a symmetric surgical instrument.

TOP-1 ADD@5 MM44.1% 57.8%

after symmetry-aware geometric candidate reranking

Rescued, harmed, and unchanged-success examples comparing ground truth, original top-1, and geometrically reranked pose projections.
What changed—and where it still failedRescued, harmed and unchanged-success cases make the method’s value and limitations visible.

Approach

A modular RGB–stereo–CAD pipeline

Segmentation and stereo depth feed Any6D candidate generation. CAD geometry and object symmetry then provide interpretable constraints for reranking.

RGB + stereo + CADSegmentation & depthAny6D candidatesGeometric reranking
970
stereo frames evaluated
968
valid Any6D outputs
86.4%
frames with a valid oracle candidate
RGB, RAFT-Stereo depth, FoundationStereo depth, and absolute depth difference.
Depth diagnosticsBest-agreement, typical and worst cases.
RGB, ground-truth mask, Mask R-CNN, and GroundingDINO plus SAM2 segmentation.
Segmentation diagnosticsRepresentative, contrasting and empty-output cases.

Selected systems

Structured AI beyond robotics

Team project · Imperial College London

Clinical multi-agent Text-to-SQL

A LangGraph workflow over MIMIC-IV that plans, generates and validates SQL, repairs failed queries, and abstains when a safe answer cannot be produced.

My contribution Designed the Planner Agent’s stepwise logic and integrated LangGraph state, nodes, transitions and failure-case testing.

17 / 20questions answered correctly in the team’s project evaluation
INPUTClinical questionNatural-language request
MY FOCUSPlanner agentIntent analysis and stepwise decompositionCohort → time window → admissions → aggregate
GENERATESQL agentSchema-grounded query construction
CHECKValidationSyntax, schema and execution checks
INVALIDRepair
VALIDExecuteGrounded answer
UNRESOLVEDSafe abstention
TNT recipe to in-game verification25-second team demo · captions available

Team project · Real-time voice system

Minecraft voice assistant grounded in live game state

A complete interaction loop combining speech input, structured game-state parsing, an LLM/API response, and parallel TTS and Gradio output.

My work: real-time state capture and parsing, structured-state validation, API integration, and controlled synchronisation and responsiveness testing.

About

Research through systems that have to work.

I work between learning-based robotics and the engineering required to make models run on physical systems. My interests include imitation learning, reinforcement learning, Sim2Real, VLA models, 3D perception and deployment-oriented evaluation.

Robot learning Teleoperation · demonstrations · VLA deployment · physical evaluation

Perception 6-DoF pose · stereo depth · segmentation · CAD geometry

Engineering Python · PyTorch · Linux · Git · hardware/software debugging

Contact

Open to PhD, research and embodied-AI opportunities.