I am a Fellow of CAE, IEEE, IAPR, and AAIA, and an ACM Distinguished Member. Prior to joining Baidu, I was a Senior Principal Researcher at Microsoft Research Asia.
My current research interests lie in world models, multimodal understanding and generation, and diffusion models. My work spans neural architecture design, semantic segmentation, and large‑scale vector similarity search.
I am best known for several widely adopted open‑source technologies and foundational algorithms: HRNet — a foundational visual backbone for high‑resolution representation learning, broadly adopted in human pose estimation, semantic segmentation, autonomous driving, and many other vision tasks, OCRNet — a transformer‑based semantic segmentation approach that improves pixel‑level labeling via object queries, and NGS/SPTAG — a production‑ready graph‑based vector search system deployed at the hundred‑billion‑data scale.
More recently, I have contributed extensively to modern generative AI and self-supervised representation learning. My recent innovations include SRA and MixFlow for diffusion training, the Hallo series for human video generation, and CAE for self-supervised visual representation learning. The core principle of SRA is extended by Self-Flow, the key training technique for the FLUX 3 generative model.
I served as Program Chair of ICCV 2025. I frequently serve as a (Senior) Area Chair for major conferences including CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, ACM MM, IJCAI, and AAAI. I serve as an Associate Editor for IEEE TPAMI, IJCV, and ACM TOMM, and previously served for IEEE TCSVT and IEEE TMM.