Skip to main navigation Skip to search Skip to main content

TransMask-Anim: Transformer-Driven Mask Perturbation and Keypoint Correspondence for Latent Space Image Animation

  • Nega Asebe Teka
  • , Zhenting Zhou
  • , Jianwen Chen
  • University of Electronic Science and Technology of China

Research output: Chapter in Book / Conference PaperConference Paperpeer-review

Abstract

Image animation aims to synthesize realistic video sequences by transferring motion from a driving video to a static source image, while maintaining source identity and temporal coherence. Challenges such as occlusion handling, identity preservation, and nonlinear motion navigation persist in existing methods. We present TransMask-Anim, a transformer-driven framework that leverages Vision Transformers (ViT) for enhanced mask perturbation and keypoint correspondence in latent space. Our innovations include: (1) ViT-Enhanced Keypoint-Mask Fusion with Dynamic Correspondence, where ViT's multi-head attention treats keypoints as query tokens and mask patches as keys/values, dynamically aligning them across frames to create fused tokens that capture long-range dependencies; perturbations are selectively applied to high-attention regions, preserving motion and masking identity, (2) Hierarchical Latent Animation with ViT-Bottleneck Navigation, using ViT encoders to generate multi-level latents (coarse for global motion, fine for details), navigated by a transformer router selecting paths based on keypoint scores, incorporating perturbations as latent noise for artifact reduction and cycle-consistent generalization to complex sequences. Self-supervised training ensures robust performance. Evaluations on VoxCeleb, TaiChiHD, and TED-Talks datasets show TransMask-Anim outperforming state-of-the-art in L1, AKD, MKR, AED, LPIPS, and FVD metrics, advancing applications in digital media and virtual reality.

Original languageEnglish
Title of host publication2025 22nd International Computer Conference on Wavelet Active Media Technology and Information Processing, ICCWAMTIP 2025
EditorsJian Ping Li, Igor Bloshanskii, Ishfaq Ahmad, Simon X. Yang, Xin Shuo Li, Hao Yu Xie
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331593766
DOIs
Publication statusPublished - 2025
Externally publishedYes
Event2025 22nd International Computer Conference on Wavelet Active Media Technology and Information Processing, ICCWAMTIP 2025 - Chengdu, China
Duration: 19 Dec 202521 Dec 2025

Conference

Conference2025 22nd International Computer Conference on Wavelet Active Media Technology and Information Processing, ICCWAMTIP 2025
Country/TerritoryChina
CityChengdu
Period19/12/2521/12/25

Bibliographical note

Publisher Copyright:
© 2025 IEEE.

Keywords

  • Image animation
  • Keypointmask fusion
  • Latent navigation
  • Temporal coherence
  • Vision Transformer

Fingerprint

Dive into the research topics of 'TransMask-Anim: Transformer-Driven Mask Perturbation and Keypoint Correspondence for Latent Space Image Animation'. Together they form a unique fingerprint.

Cite this