SyncAnimation: A Real-Time End-to-End Framework for Audio-Driven Human Pose and Talking Head Animation

* Corresponding author # Co-first author (equal contribution)

Publication
Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence (IJCAI-25), 1657–1665 (2025)

Read the conference record for the abstract and access options.

The framework uses reference-image parameters and audio to guide pose and expression, with successive stages for upper-body rendering, head rendering and lip refinement.

SyncAnimation: audio-conditioned pose, expression and progressive avatar rendering.

Figure 2 · Yujian Liu et al., SyncAnimation, arXiv:2501.14646v1 (2025). Full figure reproduced; display size and image format only are adjusted. · Authors’ paper figure · CC BY 4.0

View full-size figure