I am a final-year Ph.D. student at The Chinese University of Hong Kong, Shenzhen, supervised by Prof. Xiang Wan and Prof. Guanbin Li. I am also a research intern at AMU, Baidu. Before that, I received my Master’s degree from Sun Yat-sen University, where I worked in the HCP Laboratory under the supervision of Prof. Guanbin Li.

My long-term research vision is to build intelligent multimodal systems that can perceive, reason, and act in complex real-world environments. Toward this goal, my current research centers on two complementary directions:

1) Learning to Reason: Vision-Language Models (VLM) as Multimodal Thinkers. I study how vision-language models can develop robust reasoning and decision-making abilities, with a particular focus on agentic thinking, latent reasoning, reinforcement learning, and on-policy distillation.

2) Learning to Act: Vision-Language-Action (VLA) Systems in 3D Environments. Building on reasoning-centric vision-language models, I explore how multimodal systems can extend from visual understanding to 3D-aware perception, spatial reasoning, and action-oriented intelligence in physical environments.

I am currently on the job market and actively looking for opportunities in both industry and academia. Please feel free to contact me if you are interested in my work or potential collaboration.