Mi Yan 严汨

I am a PhD student at Peking University, advised by Prof. He Wang. I received my bachelor's degree from Turing Class, Peking University, in 2023.

I see generalization as the central challenge in embodied AI, and study how data and model design can help embodied models generalize beyond their training experience; my earlier research focused on generalization in 3D perception.

After exploring data, VLAs, and world models, I have begun to doubt whether the embodied AI community can develop intelligence on its own, and to suspect that we may instead need to rely on the intelligence of large language models. From this perspective, the role of embodied systems may be to serve as reliable trackers that translate LLM outputs into physical action, rather than develop intelligence of their own.

Portrait of Mi Yan

Exploring VLA generalization across dimensions through simulation

Five dimensions.
One generalization goal.

GraspVLA
CoRL

Our first simulation data-generation and evaluation pipeline produces an open billion-frame dataset of interactions with 10,000 objects, enabling zero-shot sim-to-real grasping without real-world fine-tuning.

Building on our simulation pipeline, StereoVLA extracts rich geometry from large-scale synthetic stereo-action data through GeoSem and 2 synergistic co-training tasks, enabling robust manipulation across near-hemispherical camera viewpoints.

GIF
CoRL

From a task instruction, GIF automatically generates interactive, functional scenes and robot trajectories, synthesizing data to train and evaluate VLA policies that transfer to real robots.

ZETA
CoRL

Using a simulated dataset spanning 512 embodiments, ZETA systematically analyzes cross-embodiment VLA transfer across state-action representations, embodiment diversity, co-training tasks, and zero-shot definitions, distinguishing strict transfer from transfer with target-embodiment pretraining exposure.

Publications

* Equal contribution

Simulation for VLA generalization

Scene & object, camera, task, and embodiment generalization.

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data — overview

Scene & object generalization

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data

Shengliang Deng*, Mi Yan*, Songlin Wei, Haixin Ma, Yuxin Yang, Jiayi Chen, Zhiqi Zhang, Taoyu Yang, Xuheng Zhang, Wenhao Zhang, Heming Cui, Zhizheng Zhang, He Wang

Our first simulation data-generation and evaluation pipeline produces an open billion-frame dataset of interactions with 10,000 objects, enabling zero-shot sim-to-real grasping without real-world fine-tuning.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision — overview

Camera generalization

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Shengliang Deng*, Mi Yan*, Yixin Zheng*, Jiayi Su, Wenhao Zhang, Xiaoguang Zhao, Heming Cui, Zhizheng Zhang, He Wang

Building on our simulation pipeline, StereoVLA extracts rich geometry from large-scale synthetic stereo-action data through GeoSem and 2 synergistic co-training tasks, enabling robust manipulation across near-hemispherical camera viewpoints.

GIF: Agentic Generation of Interactive and Functional Object Compositions for Robot Learning — overview

Task generalization

GIF: Agentic Generation of Interactive and Functional Object Compositions for Robot Learning

Long Xu*, Zhiqi Zhang*, Mi Yan*, S. Deng, C. Xia, M. Dong, J. Chen, J. Lyu, F. Gao, Z. Zhang, H. Wang

From a task instruction, GIF automatically generates interactive, functional scenes and robot trajectories, synthesizing data to train and evaluate VLA policies that transfer to real robots.

CoRL 2026
ZETA: A Controlled Study of Zero-Shot Cross-Embodiment VLA Transfer for Tabletop Manipulation — overview

Embodiment generalization

ZETA: A Controlled Study of Zero-Shot Cross-Embodiment VLA Transfer for Tabletop Manipulation

Mi Yan*, Wenhao Zhang*, Zhiqi Zhang*, Yu Peng*, Tangxinyu Wang*, Lingfei Zhai, Jiayi Su, Shengliang Deng, Lin Peng, Yaowei Liu, Yuxing Chen, Zhiyuan Wei, Jilong Wang, Jiayi Chen, Jiangran Lyu, Zhizheng Zhang, He Wang

Using a simulated dataset spanning 512 embodiments, ZETA systematically analyzes cross-embodiment VLA transfer across state-action representations, embodiment diversity, co-training tasks, and zero-shot definitions, distinguishing strict transfer from transfer with target-embodiment pretraining exposure.

Generalist robot learning

Learning reusable dynamics, robust spatial representations, and dexterous behaviors.

LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion — overview

LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion

Jiangran Lyu*, Kai Liu*, Xuheng Zhang*, Haoran Liao, Yusen Feng, Wenxuan Zhu, Tingrui Shen, Jiayi Chen, Jiazhao Zhang, Yifei Dong, Wenbo Cui, Senmao Qi, Shuo Wang, Yixin Zheng, Mi Yan, Xuesong Shi, Haoran Li, Dongbin Zhao, Ming-Yu Liu, Zhizheng Zhang, Li Yi, Yizhou Wang, He Wang

Unifies dynamics, visual prediction, and policy learning from 30,000+ hours of heterogeneous human and robot interaction data.

Emerging Extrinsic Dexterity in Cluttered Scenes via Dynamics-aware Policy Learning — overview

Emerging Extrinsic Dexterity in Cluttered Scenes via Dynamics-aware Policy Learning

Yixin Zheng*, Jiangran Lyu*, Yifan Zhang, Jiayi Chen, Mi Yan, Yuntian Deng, Xuesong Shi, Xiaoguang Zhao, Yizhou Wang, Zhizheng Zhang, He Wang

Learns contact dynamics to guide reinforcement learning, enabling robots to exploit environmental contact for manipulation in cluttered scenes.

KPGrasp: Scalable Keypoint Flow Matching for Dexterous Grasp Generation — overview

KPGrasp: Scalable Keypoint Flow Matching for Dexterous Grasp Generation

Yuansen Huang*, Jiayi Chen*, Haoran Liu, Yubin Ke, Bing Han, Jiangran Lyu, Mi Yan, Li Yi, He Wang

Represents dexterous grasps as 3D hand keypoints and learns a scalable flow-matching prior without contact losses or contact-based test-time refinement.

CoRL 2026
Adaptive Zone-aware Hierarchical Planner for Vision-Language Navigation — overview

Adaptive Zone-aware Hierarchical Planner for Vision-Language Navigation

Chen Gao, Xingyu Peng, Mi Yan, He Wang, Lirong Yang, Haibing Ren, Hongsheng Li, Si Liu

Decomposes instruction-following navigation into adaptive zone-level subgoal planning and low-level execution for long-horizon navigation.

3D perception & transfer

Understanding objects and interactions across views and domains.

MaskClustering: View Consensus based Mask Graph Clustering for Open-Vocabulary 3D Instance Segmentation

Mi Yan, Jiazhao Zhang, Yan Zhu, He Wang

Clusters 2D masks using multi-view consensus to recover open-vocabulary 3D instances without training.

Experience

Peking University emblem

Peking University

PhD studentSep. 2023 – present

Undergraduate studentSep. 2019 – Jun. 2023

Research advisor: Prof. He Wang

Academic Services

Awards

  • 2023: Peking University Presidential Scholarship (around 3 awardees a year)
  • 2023: Peking University Outstanding Graduates (Top 5%)
  • 2022: SenseTime Scholarship (around 30 awardees a year)
  • 2022: Peking University Triple-A student
  • 2021: Peking University Research Excellence Award
  • 2020: Google Women Techmaker Scholarship (around 30 awardees a year)
  • 2020: Peking University Triple-A student
  • 2019: Top ten leading student in Hangzhou Xuejun High school
  • 2017 & 2018: First prize in National Mathematical Olympiad, Zhejiang Province

Beyond Research

Outside of research, I enjoy cooking, playing sports, hiking, going for walks, spending time with friends, and playing the piano. I love watching the rain and noticing the little changes around me as the seasons turn: new leaves on the tree outside my window, buds on the roses growing by the wall, or a new dish on the menu at the little restaurant downstairs.

Wisdom, beauty, and love are at the heart of the life I hope to lead. I want to keep learning, appreciate and create beauty, and nurture meaningful connections with others.