Ziyang Song Multimodal AI Researcher, Tencent

I am a researcher at Tencent Video, working on multimodal video generation & world models for filmmaking.

I got my PhD degree (2021 ~ 2025) from The Hong Kong Polytechnic University, fortunately supervised by Prof. Bo Yang. My PhD research focused on 3D reconstruction and scene understanding. Prior to that, I got my M.Eng and B.Eng degrees (Honors Youth Program) from Xi'an Jiaotong University.

During my PhD study, I interned at TikTok (San Jose, CA) with Xinyu Gong. During my M.Eng study, I interned at SenseTime with Dongliang Wang, and Tencent Robotics X with Wanchao Chi.

profile photo
Featured Publications
NEW!
MV-S2V: Multi-View Subject-Consistent Video Generation
Ziyang Song, Xinyu Gong, Bangya Liu, Zelin Zhao
ACM SIGGRAPH, 2026
arXiv Project Page Code
NEW!
CETCAM: Camera-Controllable Video Generation via Consistent and Extensible Tokenization
Zelin Zhao, Xinyu Gong, Bangya Liu, Ziyang Song, Jun Zhang, Suhui Wu, Yongxin Chen, Hao Zhang
Computer Vision and Pattern Recognition Findings (CVPR Findings), 2026
arXiv Project Page
TRACE: Learning 3D Gaussian Physical Dynamics from Multi-view Videos
Jinxi Li, Ziyang Song, Bo Yang
International Conference on Computer Vision (ICCV), 2025
arXiv Code
FreeGave: 3D Physics Learning from Dynamic Videos by Gaussian Velocity
Jinxi Li, Ziyang Song, Siyuan Zhou, Bo Yang
Computer Vision and Pattern Recognition (CVPR), 2025
arXiv Code
OSN: Infinite Representations of Dynamic 3D Scenes from Monocular Videos
Ziyang Song, Jinxi Li, Bo Yang
International Conference on Machine Learning (ICML), 2024
arXiv Video Code
NVFi: Neural Velocity Fields for 3D Physics Learning from Dynamic Videos
Jinxi Li, Ziyang Song, Bo Yang
Advances in Neural Information Processing Systems (NeurIPS), 2023
arXiv Code
ActFormer: A GAN-based Transformer towards General Action-Conditioned 3D Human Motion Generation
Liang Xu*, Ziyang Song*, Dongliang Wang, Jing Su, Zhicheng Fang, Chenjing Ding, Weihao Gan, Yichao Yan, Xin Jin, Xiaokang Yang, Wenjun Zeng, Wei Wu
International Conference on Computer Vision (ICCV), 2023
arXiv Project Page Code

(* denotes equal contribution)
OGC: Unsupervised 3D Object Segmentation from Rigid Dynamics of Point Clouds
Ziyang Song, Bo Yang
Advances in Neural Information Processing Systems (NeurIPS), 2022
arXiv Video Code