👋 About Me
I am a Senior Algorithm Expert on the ATH-Token Foundry multimodal generation team at Alibaba Group, where I work on foundation models for video generation. I received my Ph.D. from The Hong Kong Polytechnic University under the supervision of Prof. Lei Zhang.
My research interests lie at the intersection of multimodal understanding and generation. In particular, I focus on omni-reference generation, in-context visual generation, and interactive generation. My work has been published at leading international conferences and journals, including CVPR, ICCV, ECCV, NeurIPS, ICLR, and IJCV.
I am always interested in discussing research ideas and potential collaborations. Please feel free to contact me at cssjcai@gmail.com.
🔥 News
- 2026.06: 🎉 EchoStyle has been accepted to ECCV 2026.
- 2026.02: 🎉 AnyID has been accepted to CVPR 2026.
- 2026.01: 🎉 EchoMotion has been accepted to ICLR 2026.
- 2025.09: 🎉 EchoShot has been accepted to NeurIPS 2025.
- 2025.06: 🎉 PerLDiff has been accepted to ICCV 2025.
- 2025.03: 🎉 CT3D++ has been accepted for publication in IJCV.
- 2025.01: 🎉 TAU-106k has been accepted to ICLR 2025.
📝 Publications
2026
EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis
Huaqiu Li, Jiahao Wang, Sijia Cai†, Hualian Sheng, Bing Deng, Jieping Ye, Wenhan Luo
† Project Lead
ECCV 2026
arXiv · Project Page · Code
A scalable text-driven framework for high-fidelity and long-form video stylization using reverse data synthesis.
AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual References
Jiahao Wang, Hualian Sheng, Sijia Cai†, Yuxiao Yang, Weizhan Zhang, Caixia Yan, Bing Deng, Jieping Ye
† Project Lead
CVPR 2026
arXiv · Project Page · Code: Coming Soon
An ultra-fidelity identity-preserving video generation framework supporting diverse and free-form visual references.
EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
Yuxiao Yang, Hualian Sheng, Sijia Cai†, Jing Lin, Jiahao Wang, Bing Deng, Junzhe Lu, Haoqian Wang, Jieping Ye
† Project Lead
ICLR 2026
arXiv · Project Page · Code
A unified dual-modality diffusion framework for jointly modeling human video appearance and explicit 3D motion.
2025
EchoShot: Multi-Shot Portrait Video Generation
Jiahao Wang, Hualian Sheng, Sijia Cai†, Weizhan Zhang, Caixia Yan, Yachuang Feng, Bing Deng, Jieping Ye
† Project Lead
NeurIPS 2025
arXiv · Project Page · Code
A scalable framework for identity-consistent and controllable multi-shot portrait video generation.
PerLDiff: Controllable Street View Synthesis Using Perspective-Layout Diffusion Model
Jinhua Zhang, Hualian Sheng, Sijia Cai, Bing Deng, Qiao Liang, Wen Li, Ying Fu, Jieping Ye, Shuhang Gu
ICCV 2025
A perspective-layout diffusion model that uses 3D geometric priors for precise and controllable street-view synthesis.
CT3D++: Improving 3D Object Detection with Keypoint-Induced Channel-wise Transformer
Hualian Sheng, Sijia Cai, Na Zhao, Bing Deng, Qiao Liang, Min-Jian Zhao, Jieping Ye
International Journal of Computer Vision (IJCV), 2025
A flexible 3D object detection framework based on keypoint-induced channel-wise Transformer refinement.
TAU-106K: A New Dataset for Comprehensive Understanding of Traffic Accident
Yixuan Zhou, Long Bai, Sijia Cai†, Bing Deng, Xing Xu, Heng Tao Shen
† Project Lead
ICLR 2025
[Paper](https://openreview.net/forum?id=Fb0q2uI4Ha · Code & Dataset
A large-scale multimodal dataset and specialized model for comprehensive traffic-accident understanding.