I am a final-year Ph.D. student at The Chinese University of Hong Kong, Shenzhen, supervised by Prof. Xiang Wan and Prof. Guanbin Li. I am also a research intern at AMU, Baidu. Before that, I received my Master’s degree from Sun Yat-sen University, where I worked in the HCP Laboratory under the supervision of Prof. Guanbin Li.

My long-term research vision is to build intelligent multimodal systems that can perceive, reason, and act in complex real-world environments. Toward this goal, my current research centers on two complementary directions:

1) Learning to Reason: Vision-Language Models (VLM) as Multimodal Thinkers. I study how vision-language models can develop robust reasoning and decision-making abilities, with a particular focus on agentic thinking, latent reasoning, reinforcement learning, and on-policy distillation.

2) Learning to Act: Vision-Language-Action (VLA) Systems in 3D Environments. Building on reasoning-centric vision-language models, I explore how multimodal systems can extend from visual understanding to 3D-aware perception, spatial reasoning, and action-oriented intelligence in physical environments.

I am currently on the job market and actively looking for opportunities in both industry and academia. Please feel free to contact me if you are interested in my work or potential collaboration.

πŸ”₯ News

  • 2026.06: One paper is accepted by ECCV 2026! πŸŽ‰πŸŽ‰
  • 2026.02: Two papers are accepted by CVPR 2026! πŸŽ‰πŸŽ‰
  • 2025.10: Our new survey paper on agentic MLLMs is released! πŸŽ‰πŸŽ‰
  • 2025.09: One paper is accepted by NeurIPS 2025! πŸŽ‰πŸŽ‰
  • 2025.06: Two papers are accepted by ICCV 2025! πŸŽ‰πŸŽ‰
  • 2025.03: Our survey paper on MLLM agents is accepted by Visual Intelligence! πŸŽ‰πŸŽ‰
  • 2024.09: One paper is accepted by NeurIPS 2024 as Spotlight! πŸŽ‰πŸŽ‰
  • 2024.06: One paper is accepted by ECCV 2024! πŸŽ‰πŸŽ‰

πŸ“ Publications

ECCV 2026
sym

StreamSpatial: A Benchmark and Framework for Streaming 3D Visual-Spatial Reasoning
Junlin Xie, Keyang Zhong, Quanlong Zheng, Ruifei Zhang, Kuo Wang, Yanhao Zhang, Haonan Lu, Xiang Wan, Guanbin Li
ECCV 2026

CVPR 2026
sym

EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic Retrieval
Jiashi Lin, Changhong Jiang, Xiangru Lin, Ruifei Zhang, Xinyi Zhu, Jiyao Liu, Cheng Tang, Ye Du, Shujian Gao, Junzhi Ning, Lihao Liu, Ziyan Huang, Tianbin Li, Jin Ye, Junjun He
CVPR 2026

CVPR 2026
sym

StreamRAG: Enhancing Real-Time Video Understanding with Retrieval Augmentation
Junlin Xie, Quanlong Zheng, Ruifei Zhang, Yanhao Zhang, Xiang Wan, Guanbin Li
CVPR 2026

arXiv 2025
sym

A survey on agentic multimodal large language models
Huanjin Yao†, Ruifei Zhang† (Equal Contribution), Jiaxing Huang, Jingyi Zhang, Yibo Wang, Bo Fang, Ruolin Zhu, Yongcheng Jing, Shunyu Liu, Guanbin Li, Dacheng Tao
arXiv 2025

NeurIPS 2025
sym

Intermediate Domain Alignment and Morphology Analogy for Patent-Product Image Retrieval
Haifan Gong, Xuanye Zhang, Ruifei Zhang, Yun Su, Zhuo Li, Yuhao Du, Anningzhe Gao, Xiang Wan, Haofeng Li
NeurIPS 2025

ICCV 2025
sym

AdaDrive: Self-Adaptive Slow-Fast System for Language-Grounded Autonomous Driving
Ruifei Zhang, Junlin Xie, Wei Zhang, Weikai Chen, Xiao Tan, Xiang Wan, Guanbin Li
ICCV 2025

ICCV 2025
sym

VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving
Ruifei Zhang, Wei Zhang, Xiao Tan, Sibei Yang, Xiang Wan, Xiaonan Luo, Guanbin Li
ICCV 2025

Visual Intelligence 2025
sym

Large Multimodal Agents: A Survey
Junlin Xie, Zhihong Chen, Ruifei Zhang, Guanbin Li
Visual Intelligence 2025

NeurIPS 2024 Spotlight
sym

WhodunitBench: Evaluating Large Multimodal Agents via Murder Mystery Games
Junlin Xie†, Ruifei Zhang† (Equal Contribution), Quanlong Zheng, Yanhao Zhang, Xiang Wan, Guanbin Li
NeurIPS 2024 (Spotlight)

ECCV 2024
sym

Interactive 3D Object Detection with Prompts
Ruifei Zhang, Xiangru Lin, Wei Zhang, Jincheng Lu, Xuekuan Wang, Xiao Tan, Yingying Li, Errui Ding, Jingdong Wang, Guanbin Li
ECCV 2024

CVPR 2023
sym

Advancing Visual Grounding with Scene Knowledge: Benchmark and Method
Zhihong Chen†, Ruifei Zhang† (Equal Contribution), Yibing Song, Xiang Wan, Guanbin Li
CVPR 2023

πŸ“– Educations

  • 2023.09 - Present, PhD, CUHK(Shenzhen), Shenzhen.
  • 2020.09 - 2023.06, Master, Sun Yat-sen University, Guangzhou.
  • 2016.09 - 2020.06, Bachelor, Northwest A&F University, Xi’an.

πŸŽ– Honors and Awards

  • National Scholarship x2 (2017, 2022)
  • Shenzhen Stock Exchange Scholarship (2021)
  • Outstanding Graduate Award (2020)
  • President Scholarship (2018)
  • The Third Place in AI Challenger (2018)