I'm a sencond-year PhD candidate in Computer Science at Beijing Institute for General Artificial Intelligence (BIGAI) and ShanghaiTech University. I'm fortunate to be advised by Prof. Ziyuan Jiao and Prof. Chenxi Xiao.
My current research centers on VLA and dexterous robotic hands equipped with tactile perception. I aim to enable robots to perform more complex and higher-level maneuvers.
In my previous research, I have primarily focused on human body motion, particularly interactive hand motion, which has greatly contributed to my current work. I feel incredibly fortunate to have had the guidance of Prof. Lan Xu, Prof. Jingyi Yu and Prof. Jingya Wang.
Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
Qwen Team ( co-first author & core contributor)
Website •
Arxiv •
Blog
Alignment unlocks scale, and scale unlocks generalization. Qwen-RobotManip aligns heterogeneous robot embodiments into a unified Vision-Language-Action framework for scalable multi-source robot learning. It substantially outperforms prior models across all evaluated OOD settings and ranks 1st on the RoboChallenge Table30 v1 Generalist Track, demonstrating strong generalization across unseen tasks, scenes, and robot platforms.
Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System
Qwen Team
Website •
Arxiv •
Blog
Qwen-RobotNav is a scalable navigation foundation model built for agentic navigation systems, with a parameterized interface that dynamically controls task modes and observation strategies. Trained on 15.6M samples, it achieves new state-of-the-art results across major navigation benchmarks and demonstrates strong zero-shot generalization in diverse environments.
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation
Qwen Team
Website •
Arxiv •
Blog
Qwen-RobotWorld is a language-conditioned video world model that predicts physically grounded future visual trajectories across robotic manipulation, autonomous driving, indoor navigation, and human-to-robot transfer. Built on a Double-Stream MMDiT and large-scale Embodied World Knowledge, it ranks 1st overall on EWMBench and DreamGen Bench, while demonstrating robust zero-shot generalization and multi-view consistency.
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments
Qwen Team (core contributor)
Website •
Arxiv
A unified embodied foundation model extending Qwen’s VL stack to action and trajectory generation via a DiT-based decoder, unifying manipulation, navigation, and trajectory prediction.
Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
Ye Wang*,
Pei Lin*,
Xionghui Chen*,
Haoqi Yuan,
Zhixuan Liang,
Yiyang Huang,
Anzhe Chen,
Zixing Lei,
Jie Zhang,
Tong Zhang,
Haoyang Li,
Chenxi Xiao,
Ziyuan Jiao,
Qin Jin†
Website •
Arxiv •
Code(coming)
TaF-VLA: Tactile-Force Alignment in Vision-Language-Action Models for Force-aware Manipulation
Yuzhe Huang*,
Pei Lin*,
Wanlin Li*,
Daohan Li,
Jiajun Li,
Jiaming Jiang,
Chenxi Xiao†,
Ziyuan Jiao†
Website •
Arxiv •
Code
DexMove: Learning Tactile-Guided Non-Prehensile Manipulation with Dexterous Hands
Pei Lin*,
Yuzhe Huang*,
Wanlin Li*,
Chenxi Xiao†,
Ziyuan Jiao†
The Fourteenth International Conference on Learning Representations ( ICLR 2026)
Website •
Paper
R-Tac0: A Rounded High-Frequency Transferable Monochrome Vision-based Tactile Sensor for Shape Reconstruction
Wanlin Li*,
Pei Lin*,
Meng Wang,
Chenxi Xiao,
Kaspar Althoefer,
Yao Su,
Ziyuan Jiao,
Hangxin Liu
The IEEE/RSJ International Conference on Intelligent Robots and Systems ( IROS 2025 )
Paper
PP-Tac: Paper Picking Using Omnidirectional Tactile Feedback in Dexterous Robotic Hands
Pei Lin*,
Yuzhe Huang*,
Wanlin Li*,
Jianpeng Ma,
Chenxi Xiao†,
Ziyuan Jiao†
Robotics: Science and Systems 2025 ( RSS 2025)
7th Robot Learning Workshop in ICLR 2025, oral
Website •
Paper •
ArXiv •
Code
HandDiffuse: Generative Controllers for Two-Hand Interactions via Diffusion Models
Pei Lin*,
Sihang Xu,
Hongdi Yang,
Yiran Liu,
Xin Chen,
Jingya Wang,
Jingyi Yu,
Lan Xu
Thirty-Ninth AAAI Conference on Artificial Intelligence ( AAAI 2025)
Website •
ArXiv
HumanNeRF: Efficiently Generated Human Radiance Field from Sparse Inputs
Fuqiang Zhao,
Wei Yang,
Jiakai Zhang,
Pei Lin,
Yingliang Zhang,
Jingyi Yu,
Lan Xu
Conference on Computer Vision and Pattern Recognition ( CVPR 2022)
Website •
ArXiv •
Code
Xin Suo,
Yuheng Jiang,
Pei Lin,
Yingliang Zhang,
Minye Wu,
Kaiwen Guo,
Lan Xu
Conference on Computer Vision and Pattern Recognition ( CVPR 2021)
Website •
ArXiv
Neural Free-Viewpoint Performance Rendering under Complex Human-object Interactions
Guoxing Sun,
Xin Chen,
Yizhang Chen,
Anqi Pang,
Pei Lin,
Yuheng Jiang,
Lan Xu,
Jingya Wang,
Jingyi Yu
ACM Multimedia ( ACM MM 2021 oral )
Website •
ArXiv