I am a fourth-year Ph.D. candidate in the Visual Computing Lab at The Hong Kong Polytechnic University, advised by Prof. Lei Zhang and Prof. Wenjun Zeng. Before my Ph.D., I received my M.S. from New York University and my B.S. from Central South University.
My research focuses on building reliable multimodal systems that can recognize uncertainty, adapt under distribution shifts, and reason safely in open-world environments.
- Reliable Agent: Self-Evolving Agents, Recursive Self-Improvement Agents
- Trustworthy Perception: OOD detection and test-time adaptation with VLMs and MLLMs
- MLLM Evaluation: out-of-context understanding, robustness, and trustworthy evaluation
- Medical AI: Medical MLLM, Medical Agent
| Project | Focus | Venue | Community |
|---|---|---|---|
| ANTS | Test-time MLLM understanding and adaptive negative textual spaces for OOD detection | CVPR 2026 Oral | |
| DDE | Training-free dual distribution estimation for zero-shot noisy test-time adaptation | ECCV 2026 | |
| KRNFT | Knowledge-regularized negative feature tuning for VLM-based OOD detection | ACM MM 2025 Oral | |
| MMOOC | A benchmark for out-of-context evaluation in multimodal large language models | Benchmark | |
| OpenOOD-VLM | Unified benchmarking for generalized OOD detection with vision-language models | Open Source |
-
Dual Distribution Estimation for Zero-shot Noisy Test-Time Adaptation with VLMs
ECCV 2026 · Project · Paper · Code -
ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and Reasoning
CVPR 2026 Oral · Project · Paper · Code -
Knowledge Regularized Negative Feature Tuning of Vision-Language Models
ACM MM 2025 Oral · Paper · Code -
MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
Paper · Code -
LAPT: Label-driven Automated Prompt Tuning for OOD Detection with Vision-Language Models
ECCV 2024 · Paper
GitHub language statistics reflect public repository composition, not overall proficiency.
I am interested in research collaborations around trustworthy vision-language models, open-world recognition, and test-time adaptation. For publications, code, and current projects, visit my research homepage or reach me by email.
Researching reliable multimodal intelligence for the open world.


