Yufan RenPh.D.

Staff Research Scientist at XPENG Robotics

My research focuses on spatial intelligence: how AI systems perceive, understand, and interact with the physical world. My work spans vision-language models, 3D reconstruction, and vision-language-action models, connecting visual reasoning and 3D understanding with robot learning. Previously at Meta; Ph.D. from EPFL.

Always happy to exchange ideas, discuss research, or meet for a coffee chat. Drop me an email.

Yufan Ren
Shenzhen · Shanghai

Selected Publications

All publications on Google Scholar
IronMind research overview

IronMind: Scaling Humanoid Dexterous Manipulation via Camera-Space Ego-Centric Pretraining

Huimin Pan*, Yufan Ren*†, Kunpeng Song*, Siyang Wang*, Xiwen Zhang*, Xiaoyun Hu, Zhuoxu Duan, Hanrui Zheng, Jialeng Ni, Nathan Zhao, Sibo Ma, Zhenxuan Fan, Zhongyang Che, Danny Bao, Jiacheng Wei, Jerry Bai, Xiaoyu Yue, Xiaoyang Guo, Chenyi Chen.

Technical report · arXiv 2026 · Paper · Project page

We study camera-space pretraining from human egocentric video and robot data for humanoid dexterous manipulation.

* Equal contribution. † Project lead.

VolRecon: Volume Rendering of Signed Ray Distance Functions for Generalizable Multi-View Reconstruction

VolRecon: Volume Rendering of Signed Ray Distance Functions for Generalizable Multi-View Reconstruction

Yufan Ren*, Fangjinhua Wang*, Tong Zhang, Marc Pollefeys, Sabine Süsstrunk
Conference paper · CVPR 2023, arXiv, Project Page, Code

We introduce VolRecon, a novel generalizable implicit reconstruction method with Signed Ray Distance Function (SRDF). To reconstruct the scene with fine details and little noise, VolRecon combines projection features aggregated from multi-view features, and volume features interpolated from a coarse global feature volume.

About

I am a Staff Research Scientist of Embodied AI / VLA at XPENG Robotics in Shenzhen. Previously, I was a Research Engineer at Meta Superintelligence Labs, working on post-training Llama for multimodal reasoning.

I earned my Ph.D. (2020–2025) at EPFL's Image and Visual Representation Lab, advised by Prof. Sabine Süsstrunk and Dr. Tong Zhang. My doctoral research spanned 3D reconstruction and perception, image diffusion, and large vision-language models.

During my Ph.D., I interned at NVIDIA Zurich with Dr. Alexander Millane and Meta London with Dr. Filippos Kokkinos.

I received my bachelor's degree from Zhejiang University, where I was a member of the Advanced Class of Engineering Education in Chu Kochen Honors College and received the Chu Kochen Award.

News

Introducing IronMind, our work on scaling humanoid dexterous manipulation.

Started a new role as Staff Research Scientist of Embodied AI / VLA at XPENG.

Joined XPENG in Shenzhen as Senior Engineer of Embodied AI / VLA.

Joined Meta full-time to work on multimodal reasoning.

Earlier activities
International Computer Vision Summer School 2023

International Computer Vision Summer School 2023

Università di Catania,

The school aims to provide a stimulating opportunity for young researchers and Ph.D. students. The participants will benefit from direct interaction and discussions with world leaders in Computer Vision. Participants will also have the possibility to present the results of their research, and to interact with their scientific peers, in a friendly and constructive environment.

Website chair, ICCP 2024, EPFL. Speaker at AWS LauzHack Cloud Research Day, September 2023.