Our first simulation data-generation and evaluation pipeline produces an open billion-frame dataset of interactions with 10,000 objects, enabling zero-shot sim-to-real grasping without real-world fine-tuning.
Mi Yan 严汨
I am a PhD student at Peking University, advised by Prof. He Wang. I received my bachelor's degree from Turing Class, Peking University, in 2023.
I see generalization as the central challenge in embodied AI, and study how data and model design can help embodied models generalize beyond their training experience; my earlier research focused on generalization in 3D perception.
After exploring data, VLAs, and world models, I have begun to doubt whether the embodied AI community can develop intelligence on its own, and to suspect that we may instead need to rely on the intelligence of large language models. From this perspective, the role of embodied systems may be to serve as reliable trackers that translate LLM outputs into physical action, rather than develop intelligence of their own.
Exploring VLA generalization across dimensions through simulation
Five dimensions.
One generalization goal.
Building on our simulation pipeline, StereoVLA extracts rich geometry from large-scale synthetic stereo-action data through GeoSem and 2 synergistic co-training tasks, enabling robust manipulation across near-hemispherical camera viewpoints.
From a task instruction, GIF automatically generates interactive, functional scenes and robot trajectories, synthesizing data to train and evaluate VLA policies that transfer to real robots.
Using a simulated dataset spanning 512 embodiments, ZETA systematically analyzes cross-embodiment VLA transfer across state-action representations, embodiment diversity, co-training tasks, and zero-shot definitions, distinguishing strict transfer from transfer with target-embodiment pretraining exposure.
Publications
* Equal contribution
Simulation for VLA generalization
Scene & object, camera, task, and embodiment generalization.

Scene & object generalization
GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data
Our first simulation data-generation and evaluation pipeline produces an open billion-frame dataset of interactions with 10,000 objects, enabling zero-shot sim-to-real grasping without real-world fine-tuning.

Camera generalization
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
Building on our simulation pipeline, StereoVLA extracts rich geometry from large-scale synthetic stereo-action data through GeoSem and 2 synergistic co-training tasks, enabling robust manipulation across near-hemispherical camera viewpoints.

Task generalization
GIF: Agentic Generation of Interactive and Functional Object Compositions for Robot Learning
From a task instruction, GIF automatically generates interactive, functional scenes and robot trajectories, synthesizing data to train and evaluate VLA policies that transfer to real robots.

Embodiment generalization
ZETA: A Controlled Study of Zero-Shot Cross-Embodiment VLA Transfer for Tabletop Manipulation
Using a simulated dataset spanning 512 embodiments, ZETA systematically analyzes cross-embodiment VLA transfer across state-action representations, embodiment diversity, co-training tasks, and zero-shot definitions, distinguishing strict transfer from transfer with target-embodiment pretraining exposure.
Generalist robot learning
Learning reusable dynamics, robust spatial representations, and dexterous behaviors.




KPGrasp: Scalable Keypoint Flow Matching for Dexterous Grasp Generation
Represents dexterous grasps as 3D hand keypoints and learns a scalable flow-matching prior without contact losses or contact-based test-time refinement.

3D perception & transfer
Understanding objects and interactions across views and domains.


Experience

Peking University
PhD studentSep. 2023 – present
Undergraduate studentSep. 2019 – Jun. 2023
Research advisor: Prof. He Wang
Academic Services
- Invited talk at Microsoft Research Asia & SenseTime
- TA of course Introduction to computer vision 2023 & 2024
- CoRL, NeurIPS, CVPR, AAAI Reviewer
Awards
- 2023: Peking University Presidential Scholarship (around 3 awardees a year)
- 2023: Peking University Outstanding Graduates (Top 5%)
- 2022: SenseTime Scholarship (around 30 awardees a year)
- 2022: Peking University Triple-A student
- 2021: Peking University Research Excellence Award
- 2020: Google Women Techmaker Scholarship (around 30 awardees a year)
- 2020: Peking University Triple-A student
- 2019: Top ten leading student in Hangzhou Xuejun High school
- 2017 & 2018: First prize in National Mathematical Olympiad, Zhejiang Province
Beyond Research
Outside of research, I enjoy cooking, playing sports, hiking, going for walks, spending time with friends, and playing the piano. I love watching the rain and noticing the little changes around me as the seasons turn: new leaves on the tree outside my window, buds on the roses growing by the wall, or a new dish on the menu at the little restaurant downstairs.
Wisdom, beauty, and love are at the heart of the life I hope to lead. I want to keep learning, appreciate and create beauty, and nurture meaningful connections with others.