Abstract
Image animation aims to synthesize realistic video sequences by transferring motion from a driving video to a static source image, while maintaining source identity and temporal coherence. Challenges such as occlusion handling, identity preservation, and nonlinear motion navigation persist in existing methods. We present TransMask-Anim, a transformer-driven framework that leverages Vision Transformers (ViT) for enhanced mask perturbation and keypoint correspondence in latent space. Our innovations include: (1) ViT-Enhanced Keypoint-Mask Fusion with Dynamic Correspondence, where ViT's multi-head attention treats keypoints as query tokens and mask patches as keys/values, dynamically aligning them across frames to create fused tokens that capture long-range dependencies; perturbations are selectively applied to high-attention regions, preserving motion and masking identity, (2) Hierarchical Latent Animation with ViT-Bottleneck Navigation, using ViT encoders to generate multi-level latents (coarse for global motion, fine for details), navigated by a transformer router selecting paths based on keypoint scores, incorporating perturbations as latent noise for artifact reduction and cycle-consistent generalization to complex sequences. Self-supervised training ensures robust performance. Evaluations on VoxCeleb, TaiChiHD, and TED-Talks datasets show TransMask-Anim outperforming state-of-the-art in L1, AKD, MKR, AED, LPIPS, and FVD metrics, advancing applications in digital media and virtual reality.
| Original language | English |
|---|---|
| Title of host publication | 2025 22nd International Computer Conference on Wavelet Active Media Technology and Information Processing, ICCWAMTIP 2025 |
| Editors | Jian Ping Li, Igor Bloshanskii, Ishfaq Ahmad, Simon X. Yang, Xin Shuo Li, Hao Yu Xie |
| Publisher | Institute of Electrical and Electronics Engineers Inc. |
| ISBN (Electronic) | 9798331593766 |
| DOIs | |
| Publication status | Published - 2025 |
| Externally published | Yes |
| Event | 2025 22nd International Computer Conference on Wavelet Active Media Technology and Information Processing, ICCWAMTIP 2025 - Chengdu, China Duration: 19 Dec 2025 → 21 Dec 2025 |
Conference
| Conference | 2025 22nd International Computer Conference on Wavelet Active Media Technology and Information Processing, ICCWAMTIP 2025 |
|---|---|
| Country/Territory | China |
| City | Chengdu |
| Period | 19/12/25 → 21/12/25 |
Bibliographical note
Publisher Copyright:© 2025 IEEE.
Keywords
- Image animation
- Keypointmask fusion
- Latent navigation
- Temporal coherence
- Vision Transformer
Fingerprint
Dive into the research topics of 'TransMask-Anim: Transformer-Driven Mask Perturbation and Keypoint Correspondence for Latent Space Image Animation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver