👋 About Me

I am a Senior Algorithm Expert on the ATH-Token Foundry multimodal generation team at Alibaba Group, where I work on foundation models for video generation. I received my Ph.D. from The Hong Kong Polytechnic University under the supervision of Prof. Lei Zhang.

My research interests lie at the intersection of multimodal understanding and generation. In particular, I focus on omni-reference generation, in-context visual generation, and interactive generation. My work has been published at leading international conferences and journals, including CVPR, ICCV, ECCV, NeurIPS, ICLR, and IJCV.

I am always interested in discussing research ideas and potential collaborations. Please feel free to contact me at cssjcai@gmail.com.

🔥 News

  • 2026.06: 🎉 EchoStyle has been accepted to ECCV 2026.
  • 2026.02: 🎉 AnyID has been accepted to CVPR 2026.
  • 2026.01: 🎉 EchoMotion has been accepted to ICLR 2026.
  • 2025.09: 🎉 EchoShot has been accepted to NeurIPS 2025.
  • 2025.06: 🎉 PerLDiff has been accepted to ICCV 2025.
  • 2025.03: 🎉 CT3D++ has been accepted for publication in IJCV.
  • 2025.01: 🎉 TAU-106k has been accepted to ICLR 2025.

📝 Publications

2026

ECCV 2026
EchoStyle video stylization results

EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis

Huaqiu Li, Jiahao Wang, Sijia Cai, Hualian Sheng, Bing Deng, Jieping Ye, Wenhan Luo

Project Lead

ECCV 2026

arXiv · Project Page · Code

A scalable text-driven framework for high-fidelity and long-form video stylization using reverse data synthesis.

CVPR 2026
AnyID identity-preserving video generation results

AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual References

Jiahao Wang, Hualian Sheng, Sijia Cai, Yuxiao Yang, Weizhan Zhang, Caixia Yan, Bing Deng, Jieping Ye

Project Lead

CVPR 2026

arXiv · Project Page · Code: Coming Soon

An ultra-fidelity identity-preserving video generation framework supporting diverse and free-form visual references.

ICLR 2026
EchoMotion framework and generation results

EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer

Yuxiao Yang, Hualian Sheng, Sijia Cai, Jing Lin, Jiahao Wang, Bing Deng, Junzhe Lu, Haoqian Wang, Jieping Ye

Project Lead

ICLR 2026

arXiv · Project Page · Code

A unified dual-modality diffusion framework for jointly modeling human video appearance and explicit 3D motion.

2025

NeurIPS 2025
EchoShot multi-shot portrait video generation results

EchoShot: Multi-Shot Portrait Video Generation

Jiahao Wang, Hualian Sheng, Sijia Cai, Weizhan Zhang, Caixia Yan, Yachuang Feng, Bing Deng, Jieping Ye

Project Lead

NeurIPS 2025

arXiv · Project Page · Code

A scalable framework for identity-consistent and controllable multi-shot portrait video generation.

ICCV 2025
PerLDiff controllable street-view synthesis results

PerLDiff: Controllable Street View Synthesis Using Perspective-Layout Diffusion Model

Jinhua Zhang, Hualian Sheng, Sijia Cai, Bing Deng, Qiao Liang, Wen Li, Ying Fu, Jieping Ye, Shuhang Gu

ICCV 2025

arXiv · Code

A perspective-layout diffusion model that uses 3D geometric priors for precise and controllable street-view synthesis.

IJCV 2025
CT3D++ framework for 3D object detection

CT3D++: Improving 3D Object Detection with Keypoint-Induced Channel-wise Transformer

Hualian Sheng, Sijia Cai, Na Zhao, Bing Deng, Qiao Liang, Min-Jian Zhao, Jieping Ye

International Journal of Computer Vision (IJCV), 2025

arXiv · Code

A flexible 3D object detection framework based on keypoint-induced channel-wise Transformer refinement.

ICLR 2025
TAU-106K traffic accident understanding dataset

TAU-106K: A New Dataset for Comprehensive Understanding of Traffic Accident

Yixuan Zhou, Long Bai, Sijia Cai, Bing Deng, Xing Xu, Heng Tao Shen

Project Lead

ICLR 2025

[Paper](https://openreview.net/forum?id=Fb0q2uI4Ha · Code & Dataset

A large-scale multimodal dataset and specialized model for comprehensive traffic-accident understanding.