Google DeepMind · Robotics & Embodied AI

Yuxiang Zhou周宇翔

Research Engineer at Google DeepMind • Ph.D. in Computing, Imperial College London

I am a Research Engineer at Google DeepMind in London, where I work on building generalist foundation models and embodied intelligence for the physical world—including Gemini Robotics 2, Gemini Robotics 1.5, Gemini 2.5, and RoboCat.

Prior to joining DeepMind in 2018, I completed my Ph.D. in Computing at Imperial College London under the supervision of Prof. Stefanos Zafeiriou, researching statistical deformable models, high-speed 3D mesh decoding, and third-person visual imitation learning.

01 / Research Focus

Core Research Pillars

Bridging large multimodal foundation models, reinforcement learning, and 3D visual perception to enable general-purpose physical intelligence.

PILLAR 01

Generalist Robot Foundation Models

Scaling multimodal Vision-Language-Action (VLA) architectures with embodied reasoning, long-horizon planning, cross-embodiment motion transfer, and self-improving data generation loops across diverse robotic hardware.

Gemini Robotics 2 Gemini Robotics 1.5 RoboCat Embodied Reasoning Motion Transfer
PILLAR 02

RL, Visual Imitation & Sim-to-Real

Learning dexterous manipulation and contact-rich skills (such as diverse-shape RGB stacking) via manipulator-independent third-person imitation, offline RL with policy kickstarting, and self-supervised sim-to-real adaptation.

Third-Person Imitation Offline RL Sim-to-Real RGB-Stacking Tactile Learning
PILLAR 03

3D Computer Vision & Deformable Models

Statistical deformable shape modelling, convolutional mesh decoders capable of >2,500 FPS 3D face reconstruction, adversarial UV map completion, and dense 3D pose estimation in unconstrained environments.

Convolutional Mesh Decoders 3D Morphable Models MeshGAN UV-GAN Dense Correspondence
02 / Background

Experience & Education

Over a decade of research and engineering across Google DeepMind, Google Brain, and Imperial College London.

Research & Industry Appointments

Research Engineer

Google DeepMind · London, UK
Dec 2018 – Present

Core contributor to flagship robotics foundation models (Gemini Robotics 2, Gemini Robotics 1.5, Gemini Robotics, Gemini 2.5, and RoboCat). Researching cross-embodiment manipulation, visual imitation learning, and large-scale reinforcement learning for real-world robots.

Research Intern

Google Brain & DeepMind · London, UK
Sep 2017 – Mar 2018

Joint research internship across Google Brain and DeepMind Robotics, focusing on deep reinforcement learning, third-person visual imitation learning, and robotic manipulation.

Graduate Teaching Assistant

Imperial College London · Dept. of Computing
2012 – 2017

Led small-group tutorials, laboratory sessions, and coursework assessment for undergraduate computing courses covering Haskell, C++, Java, Compilers, and the Pintos Operating System kernel.

Software Engineering Industrial Placement

Bloomberg L.P. · London, UK
Apr 2013 – Sep 2013

Member of the Bloomberg Mobile Professional team, building and shipping real-time financial analytics workflows for Bloomberg Anywhere across hybrid web and native mobile platforms.

Academic Education

Ph.D. in Computing

Imperial College London
2014 – 2019

Intelligent Behaviour Understanding Group (iBUG), supervised by Prof. Stefanos Zafeiriou.
Thesis: Dense Pose Estimation of Deformable Objects — Statistical deformable models, dense correspondence estimation without landmarks, 3D mesh decoding, and imitation learning.

MEng in Computing

Imperial College London
2010 – 2014

Four-year integrated Master of Engineering in Computing.
Master's Thesis: Predictive Remote Execution for Mobile Code Offloading; group lead for Pintos OS kernel & Microsoft Research GPUVerify Eclipse IDE plugin.

Technical Stack & Domains

Core Languages, Frameworks & Hardware
Python JAX / Haiku / Flax PyTorch TensorFlow C++ MuJoCo Vision-Language-Action (VLA) Bimanual & Dexterous Robots 3D Mesh Processing
03 / Research Output

Publications

Peer-reviewed conference papers, journal articles, and technical reports. Equal contribution is denoted by *.

Full Google Scholar Profile
arXiv · 2025 100+ Citations

Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer

Gemini Robotics Team, …, Yuxiang Zhou.

Introduces interleaved embodied thinking and cross-embodiment motion transfer, enabling generalist robots to reason through multi-stage physical tasks before acting.

@misc{geminiroboticsteam2025geminirobotics15pushing,
  title={Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer},
  author={Gemini Robotics Team and others and Zhou, Yuxiang},
  year={2025},
  eprint={2510.03342},
  archivePrefix={arXiv},
  primaryClass={cs.RO},
  url={https://arxiv.org/abs/2510.03342}
}
Gemini Robotics 1.5
Google DeepMind · 2025 5,290+ Citations

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Gemini Team, Google DeepMind (incl. Yuxiang Zhou).

Foundational technical report detailing the Gemini 2.5 model family, native thinking capabilities, multimodal understanding, and agentic reasoning across digital and physical domains.

@article{comanici2025gemini25,
  title={Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities},
  author={Comanici, Gheorghe and Bieber, Eric and Schaekermann, Mike and others and Zhou, Yuxiang},
  journal={arXiv preprint arXiv:2507.06261},
  year={2025}
}
GOOGLE DEEPMIND Gemini 2.5 Reasoning • Multimodality • Agentic Control
Google DeepMind · 2025 540+ Citations

Gemini Robotics: Bringing AI into the Physical World

Gemini Robotics Team, …, Yuxiang Zhou.

Adapts Gemini's multimodal world understanding into physical actions, achieving zero-shot and few-shot dexterous manipulation across bimanual and humanoid robot platforms.

@article{team2025gemini,
  title={Gemini Robotics: Bringing AI into the Physical World},
  author={Team, Gemini Robotics and Abeyruwan, Saminda and Ainslie, Joshua and Alayrac, Jean-Baptiste and Arenas, Montserrat Gonzalez and others and Zhou, Yuxiang},
  journal={arXiv preprint arXiv:2503.20020},
  year={2025}
}
Gemini Robotics
TMLR · 2023 330+ Citations *Equal Contribution

RoboCat: A Self-Improving Foundation Agent for Robotic Manipulation

Konstantinos Bousmalis*, Giulia Vezzani*, Dushyant Rao*, Coline Devin*, Alex X. Lee*, Maria Bauza*, Todor Davchev*, Yuxiang Zhou*, Agrim Gupta*, et al.

The first visual goal-conditioned foundation agent capable of controlling multiple robot embodiments and autonomously generating new training data to self-improve on novel tasks from as few as 100 demonstrations.

@article{bousmalis2023robocat,
  title={RoboCat: A Self-Improving Foundation Agent for Robotic Manipulation},
  author={Bousmalis, Konstantinos and Vezzani, Giulia and Rao, Dushyant and Devin, Coline and Lee, Alex X and Bauza, Maria and Davchev, Todor and Zhou, Yuxiang and Gupta, Agrim and Raju, Akhil and others},
  journal={Transactions on Machine Learning Research (TMLR)},
  year={2023}
}
RoboCat
ICLR · 2023 69 Citations

Lossless Adaptation of Pretrained Vision Models for Robotic Manipulation

Mohit Sharma, Claudio Fantacci, Yuxiang Zhou, Skanda Koppula, Nicolas Heess, Jon Scholz, Yusuf Aytar.

Demonstrates how side-network and adapter architectures can inject task-specific visuomotor control features into frozen pretrained vision backbones without catastrophic forgetting.

@inproceedings{sharma2023lossless,
  title={Lossless Adaptation of Pretrained Vision Models For Robotic Manipulation},
  author={Sharma, Mohit and Fantacci, Claudio and Zhou, Yuxiang and Koppula, Skanda and Heess, Nicolas and Scholz, Jon and Aytar, Yusuf},
  booktitle={International Conference on Learning Representations (ICLR)},
  year={2023}
}
IROS · 2022 24 Citations

How to Spend Your Robot Time: Bridging Kickstarting and Offline Reinforcement Learning for Vision-Based Robotic Manipulation

Alex X. Lee*, Coline Devin*, Jost Tobias Springenberg*, Yuxiang Zhou, Thomas Lampe, Abbas Abdolmaleki, Konstantinos Bousmalis.

@inproceedings{lee2022spend,
  title={How to Spend Your Robot Time: Bridging Kickstarting and Offline Reinforcement Learning for Vision-based Robotic Manipulation},
  author={Lee, Alex X and Devin, Coline and Springenberg, Jost Tobias and Zhou, Yuxiang and Lampe, Thomas and Abdolmaleki, Abbas and Bousmalis, Konstantinos},
  booktitle={IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)},
  year={2022}
}
How to Spend Your Robot Time
CoRL · 2021 146 Citations *Equal Contribution

Beyond Pick-and-Place: Tackling Robotic Stacking of Diverse Shapes

Alex X. Lee*, Coline Devin*, Yuxiang Zhou*, Thomas Lampe, Konstantinos Bousmalis, Jost Tobias Springenberg, Arunkumar Byravan, Abbas Abdolmaleki, et al.

Introduces the RGB-Stacking benchmark and trains a vision-based generalist policy via sim-to-real RL and offline data distillation to balance and stack geometrically complex objects.

@inproceedings{lee2021beyond,
  title={Beyond Pick-and-Place: Tackling Robotic Stacking of Diverse Shapes},
  author={Lee, Alex X and Devin, Coline Manon and Zhou, Yuxiang and Lampe, Thomas and Bousmalis, Konstantinos and Springenberg, Jost Tobias and Byravan, Arunkumar and Abdolmaleki, Abbas and Gileadi, Nimrod and Khosid, David and others},
  booktitle={5th Annual Conference on Robot Learning (CoRL)},
  year={2021}
}
RGB Stacking
RSS · 2021 26 Citations First Author

Manipulator-Independent Representations for Visual Imitation

Yuxiang Zhou, Yusuf Aytar, Konstantinos Bousmalis.

Learns disentangled, embodiment-agnostic visual state representations that allow a robot to imitate skills directly from third-person demonstrations performed by humans or different robot arms.

@inproceedings{zhou2021manipulator,
  title={Manipulator-Independent Representations for Visual Imitation},
  author={Zhou, Yuxiang and Aytar, Yusuf and Bousmalis, Konstantinos},
  booktitle={Robotics: Science and Systems (RSS)},
  year={2021}
}
Manipulator-Independent Representations
CoRL · 2020 27 Citations

Learning Rich Touch Representations Through Cross-Modal Self-Supervision

Martina Zambelli, Yusuf Aytar, Francesco Visin, Yuxiang Zhou, Raia Hadsell.

@inproceedings{zambelli2020learning,
  title={Learning rich touch representations through cross-modal self-supervision},
  author={Zambelli, Martina and Aytar, Yusuf and Visin, Francesco and Zhou, Yuxiang and Hadsell, Raia},
  booktitle={Conference on Robot Learning (CoRL)},
  pages={1415--1425},
  year={2020}
}
Rich Touch Representations
CoRL · 2020 50 Citations

Learning Dexterous Manipulation from Suboptimal Experts

Rae Jeong, Jost Tobias Springenberg, Jackie Kay, Daniel Zheng, Yuxiang Zhou, Alexandre Galashov, Nicolas Heess, Francesco Nori.

@inproceedings{jeong2020learning,
  title={Learning Dexterous Manipulation from Suboptimal Experts},
  author={Jeong, Rae and Springenberg, Jost Tobias and Kay, Jackie and Zheng, Daniel and Zhou, Yuxiang and Galashov, Alexandre and Heess, Nicolas and Nori, Francesco},
  booktitle={Conference on Robot Learning (CoRL)},
  year={2020}
}
Dexterous Manipulation
ICRA · 2020 99 Citations

Self-Supervised Sim-to-Real Adaptation for Visual Robotic Manipulation

Rae Jeong, Yusuf Aytar, David Khosid, Yuxiang Zhou, Jackie Kay, Thomas Lampe, Konstantinos Bousmalis, Francesco Nori.

@inproceedings{jeong2020self,
  title={Self-supervised sim-to-real adaptation for visual robotic manipulation},
  author={Jeong, Rae and Aytar, Yusuf and Khosid, David and Zhou, Yuxiang and Kay, Jackie and Lampe, Thomas and Bousmalis, Konstantinos and Nori, Francesco},
  booktitle={2020 IEEE International Conference on Robotics and Automation (ICRA)},
  pages={2718--2724},
  year={2020},
  organization={IEEE}
}
Sim-to-Real Adaptation
CVPR · 2020 3,450+ Citations

RetinaFace: Single-Stage Dense Face Localisation in the Wild

Jiankang Deng, Jia Guo, Yuxiang Zhou, Jinke Yu, Irene Kotsia, Stefanos Zafeiriou.

Unifies face box prediction, 2D facial landmark localization, and dense 3D mesh regression in a single-stage multi-task network; one of the most widely adopted face detectors in computer vision.

@inproceedings{deng2020retinaface,
  title={RetinaFace: Single-stage dense face localisation in the wild},
  author={Deng, Jiankang and Guo, Jia and Zhou, Yuxiang and Yu, Jinke and Kotsia, Irene and Zafeiriou, Stefanos},
  booktitle={IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  year={2020}
}
RetinaFace
CVPR · 2019 104 Citations First Author

Dense 3D Face Decoding Over 2500FPS: Joint Texture & Shape Convolutional Mesh Decoders

Yuxiang Zhou, Jiankang Deng, Irene Kotsia, Stefanos Zafeiriou.

Proposes joint graph/mesh convolutional decoders that reconstruct dense 3D facial geometry and albedo texture from a single in-the-wild image at over 2,500 frames per second.

@inproceedings{zhou2019dense,
  title={Dense 3d face decoding over 2500fps: Joint texture \& shape convolutional mesh decoders},
  author={Zhou, Yuxiang and Deng, Jiankang and Kotsia, Irene and Zafeiriou, Stefanos},
  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
  pages={1097--1106},
  year={2019}
}
Dense 3D Face Decoding
arXiv · 2019 98 Citations

MeshGAN: Non-Linear 3D Morphable Models of Faces

Shiyang Cheng, Michael Bronstein, Yuxiang Zhou, Irene Kotsia, Maja Pantic, Stefanos Zafeiriou.

@article{cheng2019meshgan,
  title={Meshgan: Non-linear 3d morphable models of faces},
  author={Cheng, Shiyang and Bronstein, Michael and Zhou, Yuxiang and Kotsia, Irene and Pantic, Maja and Zafeiriou, Stefanos},
  journal={arXiv preprint arXiv:1903.10384},
  year={2019}
}
MeshGAN
CVPR · 2018 297 Citations

UV-GAN: Adversarial Facial UV Map Completion for Pose-Invariant Face Recognition

Jiankang Deng, Shiyang Cheng, Niannan Xue, Yuxiang Zhou, Stefanos Zafeiriou.

@inproceedings{deng2018uv,
  title={Uv-gan: Adversarial facial uv map completion for pose-invariant face recognition},
  author={Deng, Jiankang and Cheng, Shiyang and Xue, Niannan and Zhou, Yuxiang and Zafeiriou, Stefanos},
  booktitle={Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition},
  pages={7093--7102},
  year={2018}
}
UV-GAN
IEEE FG · 2018 First Author

Improve Accurate Pose Alignment and Action Localization by Dense Pose Estimation

Yuxiang Zhou, Jiankang Deng, Stefanos Zafeiriou.

@inproceedings{zhou2018improve,
  title={Improve accurate pose alignment and action localization by dense pose estimation},
  author={Zhou, Yuxiang and Deng, Jiankang and Zafeiriou, Stefanos},
  booktitle={2018 13th IEEE International Conference on Automatic Face \& Gesture Recognition (FG 2018)},
  pages={480--484},
  year={2018},
  organization={IEEE}
}
Dense Pose Estimation
IEEE FG · 2018 73 Citations

Cascade Multi-View Hourglass Model for Robust 3D Face Alignment

Jiankang Deng, Yuxiang Zhou, Shiyang Cheng, Stefanos Zafeiriou.

@inproceedings{deng2018cascade,
  title={Cascade multi-view hourglass model for robust 3d face alignment},
  author={Deng, Jiankang and Zhou, Yuxiang and Cheng, Shiyang and Zaferiou, Stefanos},
  booktitle={2018 13th IEEE International Conference on Automatic Face \& Gesture Recognition (FG 2018)},
  pages={399--403},
  year={2018},
  organization={IEEE}
}
Cascade Multi-view Hourglass
IEEE TIP · 2019 134 Citations

Joint Multi-View Face Alignment in the Wild

Jiankang Deng, George Trigeorgis, Yuxiang Zhou, Stefanos Zafeiriou.

@article{deng2019joint,
  title={Joint multi-view face alignment in the wild},
  author={Deng, Jiankang and Trigeorgis, George and Zhou, Yuxiang and Zafeiriou, Stefanos},
  journal={IEEE Transactions on Image Processing},
  volume={28},
  number={7},
  pages={3636--3648},
  year={2019},
  publisher={IEEE}
}
Joint Multi-View Face Alignment
CVPRW · 2017 262 Citations

Marginal Loss for Deep Face Recognition

Jiankang Deng, Yuxiang Zhou, Stefanos Zafeiriou.

@inproceedings{deng2017marginal,
  title={Marginal loss for deep face recognition},
  author={Deng, Jiankang and Zhou, Yuxiang and Zafeiriou, Stefanos},
  booktitle={Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops},
  pages={60--68},
  year={2017}
}
Marginal Loss
IEEE FG · 2017 93 Citations First Author

Deformable Models of Ears In-the-Wild for Alignment and Recognition

Yuxiang Zhou, Stefanos Zafeiriou.

@inproceedings{zhou2017deformable,
  title={Deformable Models of Ears in-the-wild for Alignment and Recognition},
  author={Zhou, Yuxiang and Zaferiou, Stefanos},
  booktitle={2017 12th IEEE International Conference on Automatic Face \& Gesture Recognition (FG 2017)},
  pages={626--633},
  year={2017},
  organization={IEEE}
}
Deformable Ear Models
CVPR · 2016 First Author

Estimating Correspondences of Deformable Objects "In-the-Wild"

Yuxiang Zhou, Epameinondas Antonakos, Joan Alabort-i-Medina, Anastasios Roussos, Stefanos Zafeiriou.

@inproceedings{zhou2016estimating,
  title={Estimating correspondences of deformable objects 'in-the-wild'},
  author={Zhou, Yuxiang and Antonakos, Epameinondas and Alabort-i-Medina, Joan and Roussos, Anastasios and Zafeiriou, Stefanos},
  booktitle={Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition},
  pages={5791--5801},
  year={2016}
}
Deformable Object Correspondences
ICASSP · 2016

Semi-Autonomous Data Enrichment Based on Cross-Task Labelling of Missing Targets for Holistic Speech Analysis

Yue Zhang, Yuxiang Zhou, Jie Shen, Björn Schuller.

@inproceedings{zhang2016semi,
  title={Semi-autonomous data enrichment based on cross-task labelling of missing targets for holistic speech analysis},
  author={Zhang, Yue and Zhou, Yuxiang and Shen, Jie and Schuller, Bj{\"o}rn},
  booktitle={2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
  pages={6090--6094},
  year={2016},
  organization={IEEE}
}
Holistic Speech Analysis
IEEE QSIC · 2012 20 Citations

Fingertips Detection Algorithm Based on Skin Colour Filtering and Distance Transformation

Yun Liao, Yuxiang Zhou, Hua Zhou, Zhihong Liang.

@inproceedings{liao2012fingertips,
  title={Fingertips detection algorithm based on skin colour filtering and distance transformation},
  author={Liao, Yun and Zhou, Yuxiang and Zhou, Hua and Liang, Zhihong},
  booktitle={2012 12th International Conference on Quality Software},
  pages={276--281},
  year={2012},
  organization={IEEE}
}
Fingertips Detection

No publications match your current filter or search query.

BibTeX citation copied to clipboard