About

I am a Ph.D. candidate in the Department of Automation at the University of Science and Technology of China, expected to graduate in June 2027. I am advised by Associate Professor Yang Cao (曹洋) and Professor Zheng-Jun Zha (查正军). My research experience includes embodied intelligence, multimodal learning, and out-of-distribution detection.

I am currently an Applied Research Intern at Robbyant(蚂蚁灵波科技), supervised by Kecheng Zheng (郑可成), where I contribute to the research and development of vision-language-action (VLA) foundation models for embodied intelligence. My current work involves large-scale VLA models for real-world robotic manipulation, including model architecture construction, training infrastructure, systematic ablations, and large-scale pretraining.

Experience & Research Interests

Robbyant(蚂蚁灵波科技)
Applied Research Intern, Embodied AI · 2025.08 – Present

Deeply involved in the research and development of LingBot-VLA and LingBot-VLA 2.0, large-scale vision-language-action foundation models for real-world robotic manipulation. My contributions include model architecture construction, codebase infrastructure design, large-scale pretraining, and the implementation and optimization of an MoE-based action expert architecture for LingBot-VLA 2.0.

Ant Research
Research Intern · 2023.09 – 2025.08

Conducted research on multimodal learning, including CLIP pre-training with long captions, VLM training with generative visual representations, and the development of a VLM benchmark for evaluating the quality of detailed image captions.

University of Science and Technology of China
2022.09 – Expected Jun. 2027

Started my master's studies in the Department of Automation in September 2022 and transitioned to the Ph.D. stage in September 2024 through an integrated M.S.–Ph.D. program. My early research focused on out-of-distribution detection.

My broader research interests include embodied intelligence, multimodal learning, and large-scale pretraining.

Embodied IntelligenceVLA Foundation ModelsMultimodal LearningLarge-scale Pretraining

Highlighted Projects

LingBot-VLA 2.0: From Foundation to Application

Student first author.

LingBot-VLA 2.0 is a practical vision-language-action foundation model designed to move from large-scale pretraining toward reliable real-world robot applications. As a student first author, I focused on model architecture construction, codebase infrastructure design, and the implementation and optimization of an MoE-based action expert architecture.

LingBot-VLA: A Pragmatic VLA Foundation Model

Co-first author; leading student contributor.

LingBot-VLA is a large-scale vision-language-action foundation model for real-world robotic manipulation. I contributed as a co-first author and leading student contributor, focusing on model architecture construction, systematic ablations, codebase infrastructure design, and large-scale pretraining.

Publications

2026
Main figure of LingBot-VLA 2.0
LingBot-VLA 2.0: From Foundation to Application — Improving VLA Models in Practice
Wei Wu, Fangjing Wang, Fan Lu, He Sun, Shi Liu, Yunnan Wang, Yibin Yan, Yong Wang, Shuailei Ma, et al.
arXiv preprint arXiv:2607.06403, 2026
Student first author.

A practical VLA foundation model for real-world robot applications, with expanded cross-embodiment capabilities and an MoE-based action expert.

Main figure of LingBot-VLA: A Pragmatic VLA Foundation Model
LingBot-VLA: A Pragmatic VLA Foundation Model
Wei Wu*, Fan Lu*, Yunnan Wang, Shuai Yang, Shi Liu, Fangjing Wang, Qian Zhu, He Sun, Yong Wang, Shuailei Ma, et al.
arXiv preprint arXiv:2601.18692, 2026
*Equal contribution. Fan Lu is a student first author.

A large-scale vision-language-action foundation model for real-world robotic manipulation.

Main figure of Self-Consistent Latent Reasoning
Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model
Chenfeng Wang, Wei He, Xuhan Zhu, Chunpeng Zhou, Qizhen Li, Song Yan, Yufei Zheng, Chengjun Yu, Fan Lu, Wei Zhai, et al.
arXiv preprint arXiv:2605.12163, 2026
Main figure of Diffusion Guided Chain-of-Vision
Diffusion Guided Chain-of-Vision for Large Autoregressive Vision Models
Xinyang Wang, Kecheng Zheng, Minfeng Zhu, Wei Wu, Fan Lu, Wei Zhai, Wei Chen
CVPR, 2026
2025
Main figure of CompreCap
Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
Fan Lu, Wei Wu, Kecheng Zheng, Shuailei Ma, Biao Gong, Jiawei Liu, Wei Zhai, Yang Cao, Yujun Shen, Zheng-Jun Zha
CVPR, 2025

A benchmark for evaluating comprehensive image captioning in large vision-language models using directed scene graphs.

Main figure of Learning Visual Generative Priors without Text
Learning Visual Generative Priors without Text
Shuailei Ma, Kecheng Zheng, Ying Wei, Wei Wu, Fan Lu, Yifei Zhang, Chen-Wei Xie, Biao Gong, Jiapeng Zhu, Yujun Shen
CVPR, 2025
Main figure of Vision-Centric Activation and Coordination
Vision-Centric Activation and Coordination for Multimodal Large Language Models
Yunnan Wang, Fan Lu, Kecheng Zheng, Ziyuan Huang, Ziqiang Li, Wenjun Zeng, Xin Jin
arXiv preprint arXiv:2510.14349, 2025
2024
Main figure of LoTLIP
LoTLIP: Improving Language-Image Pre-training for Long Text Understanding
Wei Wu, Kecheng Zheng, Shuailei Ma, Fan Lu, Yuxin Guo, Yifei Zhang, Wei Chen, Qingpei Guo, Yujun Shen, Zheng-Jun Zha
NeurIPS, 2024
Main figure of DreamLIP
DreamLIP: Language-Image Pre-training with Long Captions
Kecheng Zheng, Yifei Zhang, Wei Wu, Fan Lu, Shuailei Ma, Xin Jin, Wei Chen, Yujun Shen
ECCV, 2024
2023
Main figure of Uncertainty-Aware Optimal Transport
Uncertainty-Aware Optimal Transport for Semantically Coherent Out-of-Distribution Detection
Fan Lu, Kai Zhu, Wei Zhai, Kecheng Zheng, Yang Cao
CVPR, 2023

An uncertainty-aware optimal transport framework for semantically coherent out-of-distribution detection.

Main figure of Likelihood-Aware Semantic Alignment
Likelihood-Aware Semantic Alignment for Full-Spectrum Out-of-Distribution Detection
Fan Lu, Kai Zhu, Kecheng Zheng, Wei Zhai, Yang Cao
arXiv preprint arXiv:2312.01732, 2023

A semantic alignment approach for full-spectrum out-of-distribution detection under both semantic and covariate shifts.

Contact

Email: lufan@mail.ustc.edu.cn / lufanzwj@gmail.com    WeChat: luDZ310711

Google Scholar   GitHub