👋 About Me

I’m a Research Scientist at Alibaba Wan Team, where I serve as a core contributor of Wan3.0, Wan2.7, Wan2.6 and Wan2.5. I also hold a research position at Zhejiang University, working closely with Prof. Yi Yang. I obtained my Ph.D. in Computer Science from Fudan University, supervised by Prof. Yu-Gang Jiang (IEEE Fellow) and Prof. Zuxuan Wu.

Research Interests

  • Alignment reinforcement learning alignment for visual generation (RLHF, preference optimization, reward modeling)
  • Generative models text-to-video generation, controllable visual generation, video editing
  • Representation learning video understanding, 3D understanding, image retrieval

🔥 News

  • 🎈Achieved 1400+ citations on Google Scholar and an h-index of 18.
  • Wan Team released Wan 3.0. Feel free to try it out!
  • DiffusionOPD accepted to SIGGRAPH ASIA 2026.
  • FlashMotion and FlashPortrait accepted to CVPR 2026.
  • Wan Team released Wan 2.6. Feel free to try it out!

📝 Publications

A full publication list is available on [Google Scholar][Semantic Scholar]

(*: equal contribution; †: project leader)

Selected Publications

Image Generation DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models

SIGGRAPH Asia 2026

DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models

Quanhao Li, Junqiu Yu, Kaixun Jiang, Yujie Wei, Zhen Xing, Pandeng Li, Ruihang Chu, Shiwei Zhang, Yu Liu, Zuxuan Wu

ACM SIGGRAPH Asia, 2026

Video Generation A Survey on Video Diffusion Models

CSUR 2025

A Survey on Video Diffusion Models

Zhen Xing, Qijun Feng, Haoran Chen, Qi Dai, Han Hu, Hang Xu, Zuxuan Wu, Yu-Gang Jiang

ACM Computing Surveys (CSUR, IF=28.0), 2025

Surveying 300+ recent literatures on video generation and editing with diffusion models. Achieving GitHub 2100+ stars.

Video Generation AID: Adapting Image2Video Diffusion Models for Instruction-based Video Prediction

ICCV 2025

AID: Adapting Image2Video Diffusion Models for Instruction-based Video Prediction

Zhen Xing, Qi Dai, Zejia Weng, Zuxuan Wu, Yu-Gang Jiang

International Conference on Computer Vision (ICCV), 2025

Video Generation MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance

ICCV 2025

MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance

Quanhao Li*, Zhen Xing*†, Rui Wang, Hui Zhang, Zuxuan Wu

International Conference on Computer Vision (ICCV), 2025

Video Generation SimDA: A Simple Diffusion Adapter for Efficient Video Generation

CVPR 2024

SimDA: A Simple Diffusion Adapter for Efficient Video Generation

Zhen Xing, Qi Dai, Han Hu, Zuxuan Wu, Yu-Gang Jiang

IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024

The first Parameter-efficient Text-to-Video generation model.

Video Generation StableAnimator: High-Quality Identity-Preserving Human Image Animation

CVPR 2025

StableAnimator: High-Quality Identity-Preserving Human Image Animation

Shuyuan Tu, Zhen Xing, Xintong Han, Zhi-Qi Cheng, Qi Dai, Chong Luo, Zuxuan Wu

IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025

Achieving GitHub 1300+ stars.

Video Understanding SVFormer: Semi-supervised Video Transformer for Action Recognition

CVPR 2023

SVFormer: Semi-supervised Video Transformer for Action Recognition

Zhen Xing, Qi Dai, Han Hu, Jingjing Chen, Zuxuan Wu, Yu-Gang Jiang

IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023

Other Publications

  • ECCV 2026

    DeRA: Decoupled Representation Alignment for Video TokenizationPengbo Guo, Junke Wang, Zhen Xing, Chengxu Liu, Daoguo Dong, Xueming Qian, Zuxuan Wu
  • CVPR 2026

    FlashMotion: Few-Step Controllable Video Generation with Trajectory GuidanceQuanhao Li, Zhen Xing, Rui Wang, Haidong Cao, Qi Dai, Daoguo Dong, Zuxuan Wu
  • CVPR 2026

    FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent PredictionShuyuan Tu, Yueming Pan, Yinming Huang, Xintong Han, Zhen Xing, Qi Dai, Kai Qiu, Chong Luo, Zuxuan Wu
  • EMNLP 2025

    ProLongVid: A Simple but Strong Baseline for Long-context Video Instruction TuningRui Wang, Bohao Li, Xiyang Dai, Jianwei Yang, Yi-Ling Chen, Zhen Xing, Yifan Yang, Dongdong Chen, Xipeng Qiu, Zuxuan Wu, Yu-Gang Jiang
  • NeurIPS 2024

    Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and AlgorithmsMiaosen Zhang, Yixuan Wei, Zhen Xing, Yifei Ma, Zuxuan Wu, Ji Li, Zheng Zhang, Qi Dai, Chong Luo, Xin Geng, Baining Guo
  • NeurIPS 2024

    GenRec: Unifying Video Generation and Recognition with Diffusion ModelsZejia Weng, Xitong Yang, Zhen Xing, Zuxuan Wu, Yu-Gang Jiang
  • ACL 2023

    TranSFormer: Slow-Fast Transformer for Machine TranslationBei Li, Yi Jing, Xu Tan, Zhen Xing, Tong Xiao, Jingbo Zhu
  • CVPR 2023

    PanoSwin: a Pano-style Swin Transformer for Panorama UnderstandingZhixin Ling, Zhen Xing, Manliang Cao, Xiangdong Zhou
  • ECCV 2022

    Semi-supervised Single-view 3D Reconstruction via Prototype Shape PriorsZhen Xing, Hengduo Li, Zuxuan Wu, Yu-Gang Jiang
  • ECCV 2022

    Few-shot Single-view 3D Reconstruction with Memory Prior Contrastive NetworkZhen Xing, Yijiang Chen, Zhixin Ling, Xiangdong Zhou, Yu Xiang
  • ECCV 2022

    Conditional Stroke Recovery for Fine-Grained Sketch-Based Image RetrievalZhixin Ling, Zhen Xing, Jian Zhou, Xiangdong Zhou

🎖 Honors and Awards

Below, I exhaustively list some of my Honors and Awards that inspire me a lot.

  • 2025 Outstanding graduates of Shanghai Top-1%, PhD
  • 2025 Alibaba Star Program and Tencent Qingyun Plan
  • 2024 Tencent academic scholarship Top-3%, PhD
  • 2023 Fudan University excellent academic scholarship Top-5%, PhD
  • 2023 "Star of Tomorrow" intern of MicroSoft Research Asia Top 10%, PhD
  • 2022 Tencent academic scholarship Rank 1/130, PhD
  • 2021 Fudan University excellent academic scholarship Top-5%, Master
  • 2020 Outstanding graduates of TianJin University Top-5%
  • 2018 Excellent monitor of Tianjin University Top-10
  • 2017-2020 Academic scholarship of Tianjin University Top-10%

💬 Invited Talks

  • 2024.04Talk at ByteDanceSimDA: Simple Diffusion Adapter for Efficient Video GenerationSlides
  • 2024.02Tutorial at OpenmmlabThe Past and Present of Video Diffusion ModelsSlides
  • 2023.12Tutorial at Kunlun ResearchThe Tutorial of Video Generative ModelsSlides

💻 Internships

  • MicroSoft Research Asia2022.03 - 2023.08
    Visual Computing Group
    • Video Diffusion Model
    • Video Understanding
    Advisor: Qi Dai, Han Hu

🎓 Academic Service

Conference Program Committee

  • 2022-2025IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
  • 2023-2025IEEE/CVF International Conference on Computer Vision (ICCV)
  • 2022-2024European Conference on Computer Vision (ECCV)
  • 2024-2025International Conference on Learning Representations (ICLR)
  • 2025ACM SIGGRAPH Conference (SIGGRAPH)
  • 2024-2025International Conference on Machine Learning (ICML)
  • 2024-2025Conference on Neural Information Processing Systems (NeurIPS)
  • 2023-2025AAAI Conference on Artificial Intelligence (AAAI)

Journal Reviewer

  • IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)
  • International Journal of Computer Vision (IJCV)
  • IEEE Transactions on Multimedia (TMM)
  • IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)
  • Knowledge-Based Systems (KBS)
  • Pattern Recognition (PR)