About Me
I am Junhao Shi, a second-year PhD student at Fudan University, SII. My research centers on large language models, agent models, and embodied intelligence, with a particular interest in how foundation models can reason, plan, and act in open-ended environments.
My recent work spans embodied task planning, robotic control, world modeling, and robust generalization. I am especially interested in building agents that can connect language understanding with grounded decision-making, and in developing training recipes that make these systems more robust, efficient, and useful in the real world.
A fuller publication record is available on Google Scholar.
News
- 2026.07: Released the arXiv preprint Learning to Move Before Learning to Do: Task-Agnostic Pretraining for VLAs.
- 2026.06: Released the arXiv preprint Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy.
- 2026.06: Released the arXiv preprint In-Context World Modeling for Robotic Control.
- 2026.05: Released the survey World Action Models: The Next Frontier in Embodied AI.
- 2025.07: Our paper How to Mitigate Overfitting in Weak-to-Strong Generalization? appeared at ACL 2025.
Selected Publications

Learning to Move Before Learning to Do: Task-Agnostic Pretraining for VLAs
Junhao Shi, Siyin Wang, Xiaopeng Yu, Li Ji, Jingjing Gong, Xipeng Qiu
- Proposes task-agnostic pretraining to improve action learning and transfer for vision-language-action models.

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy
Junhao Shi, Z. Huai, Siyin Wang, J. Chen, Y. Wang, Zhaoye Fei, H. Chen, Jingjing Gong, Xipeng Qiu, et al.
- Pushes embodied agents toward everyday physical autonomy by combining isolated skills into a unified omnimodal system.

In-Context World Modeling for Robotic Control
Siyin Wang, Junhao Shi, S. Fei, Z. Fu, L. Ji, Jingjing Gong, Xipeng Qiu
- Studies in-context world modeling as a way to improve robotic control and generalize across embodied settings.

World Action Models: The Next Frontier in Embodied AI
Siyin Wang, Junhao Shi, Z. Fu, X. He, F. Liu, C. Yang, Y. Zhou, Zhaoye Fei, Jingjing Gong, Jinlan Fu, et al.
- A survey-style overview of world action models and the emerging path from language understanding to embodied action.

LIBERO-Plus: A Progressive Robustness Benchmark for Visual-Language-Action Models
Senyu Fei, Siyin Wang, Junhao Shi, Zihao Dai, Jikun Cai, Pengfang Qian, Li Ji, Xinzhe He, Shiduo Zhang, Zhaoye Fei, Jinlan Fu, Jingjing Gong, Xipeng Qiu
- A systematic robustness study of vision-language-action models, covering failure modes and evaluation settings that matter for embodied deployment.

Yicheng Liu, Shiduo Zhang, Zibin Dong, Baijun Ye, Tianyuan Yuan, Xiaopeng Yu, Linqi Yin, Chenhao Lu, Junhao Shi, Luca Jiang-Tao Yu, Liangtao Zheng, Tao Jiang, Jingjing Gong, Xipeng Qiu, Hang Zhao
- Improves autoregressive VLA efficiency through neural action tokenization for scalable embodied modeling.

RoboOmni: Proactive Robot Manipulation in Omni-modal Context
Siyin Wang, Jinlan Fu, Feihong Liu, Xinzhe He, Huangxuan Wu, Junhao Shi, Kexin Huang, Zhaoye Fei, Jingjing Gong, Zuxuan Wu, Yu-Gang Jiang, See-Kiong Ng, Tat-Seng Chua, Xipeng Qiu
- Explores proactive robot manipulation under rich omni-modal context, bridging perception, planning, and action.

World-aware Planning Narratives Enhance Large Vision-Language Model Planner
Junhao Shi, Zhaoye Fei, Siyin Wang, Qipeng Guo, Jingjing Gong, Xipeng Qiu
- Improves embodied planning by injecting structured world-aware narratives into large vision-language model planners.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
Zhaoye Fei, Lifan Ji, Siyin Wang, Junhao Shi, Jingjing Gong, Xipeng Qiu
- Improves embodied task planning by combining LLM reasoning with reinforcement learning for grounded planning behavior.

How to Mitigate Overfitting in Weak-to-Strong Generalization?
Junhao Shi, Qinyuan Cheng, Zhaoye Fei, Yining Zheng, Qipeng Guo, Xipeng Qiu
- Studies weak-to-strong generalization in language models and explores how to reduce overfitting during transfer from weaker supervision.
Research Interests
- Large language models
- Agent models
- Embodied intelligence
Experience
- Current: PhD Student, Fudan University
- Research focus: Building foundation-model-driven agents that can understand, plan, and act in complex physical environments
Representative Topics
- Embodied task planning with large vision-language models
- Efficient training and pretraining for vision-language-action models
- World modeling for robotic control
- Robustness and generalization analysis for embodied agents