V2E Face: Emotionally Disentangled Talking Head Generation with Vector Quantization and Attention Fusion
V2E Face disentangles head pose, landmarks, and emotion using separate supervised VQ codebooks and hierarchical attention fusion, achieving SOTA lip-sync and reconstruction on CREMA-D and MEAD while staying lightweight and real-time. Published in IEEE MultiMedia, 2026.