Publication
V2E Face: Emotionally Disentangled Talking Head Generation with Vector Quantization and Attention Fusion
IEEE MultiMedia (2026) · DOI: 10.1109/MMUL.2026.3724384
Most talking-head generators tangle identity, pose, and emotion together in one latent space, then try to untangle them with extra loss terms. V2E Face takes a different approach: disentanglement is built into the architecture itself.

Core Idea#
- Multi-Attribute Vector Quantization (MAVQ) — separate, supervised VQ codebooks for yaw, pitch, roll, and facial landmarks, isolating geometry from expression from the very first bottleneck.
- Multi-head Attention Fusion (MAF) — fuses the three pose codes into one physically coherent pose vector instead of naive concatenation.
- Hierarchical attention synthesis — builds a stable geometric "canvas" first, then applies emotion as a distinct layer on top.
Everything (pose, landmarks, lip motion) is generated purely from audio + a reference image, no external pose input needed.
Results#
Evaluated on CREMA-D and MEAD against MakeItTalk, EMMN, EAT, and ED-Talk:
| Dataset | PSNR ↑ | SSIM ↑ | Sync ↑ | Acc_emo ↑ |
|---|---|---|---|---|
| CREMA-D | 25.49 (best) | 0.802 (best) | 8.79 (best) | 76.4 |
| MEAD | 23.50 (best) | 0.734 (best) | 8.32 | 67.1 |
- Best PSNR/SSIM/lip-sync on both datasets; emotion accuracy close to the top baseline.
- Human study (17 participants): ranked best among all generative methods on visual quality, motion naturalness, emotion expressiveness, and overall preference (p < 0.01 vs. baselines).
- Ablation: removing hierarchical synthesis causes a 14-point crash in emotion accuracy — confirming structured fusion, not just parameter count, drives the gain.
Ablation study (MEAD):
| Method | PSNR ↑ | SSIM ↑ | FID ↓ | Sync ↑ | Acc_emo ↑ |
|---|---|---|---|---|---|
| Proposed | 23.500 | 0.734 | 22.121 | 8.322 | 67.080 |
| w/o VQ_lmk | 18.327 | 0.594 | 24.763 | 5.211 | 66.483 |
| w/o VQ_pose | 19.882 | 0.662 | 23.302 | 5.565 | 66.805 |
| w/o MAF | 19.332 | 0.657 | 26.801 | 6.003 | 60.032 |
| w/o HS | 17.501 | 0.540 | 28.367 | 5.477 | 53.272 |
- Lightweight & fast: 15.6M parameters, 2.4 ms/frame on an RTX 4090, 24 FPS end-to-end.
Human evaluation (blind ranking, 17 participants, 1 = best, 4 = worst):
| Method | Visual Quality ↓ | Motion Naturalness ↓ | Emotion Expr. ↓ | Overall Pref. ↓ |
|---|---|---|---|---|
| GT (ref.) | 1.68 | 1.47 | 1.44 | 1.44 |
| Ours | 2.00 | 2.21 | 2.24 | 2.26 |
| EDTalk | 3.29 | 3.03 | 3.12 | 3.03 |
| MakeItTalk | 3.03 | 3.29 | 3.21 | 3.26 |

Known limitation#
Fear and Happy share an open-mouth articulation, so the model sometimes confuses them — a geometry/articulation issue rather than a disentanglement failure.

Ethics note#
As with any photorealistic talking-head generator: risk of non-consensual identity impersonation. The authors recommend pairing this kind of framework with forgery-detection and watermarking.
BibTeX#
@article{vo2026v2eface,
author = {Vo, Hoang-Son and Yang, Hyung-Jeong and Kim, Soo-Hyung},
title = {Emotionally Disentangled Talking Head Generation with Vector Quantization and Attention Fusion},
journal = {IEEE MultiMedia},
year = {2026},
doi = {10.1109/MMUL.2026.3724384}
}