Face Frontalization Using Vision Transformers Görsel Dönüştürücüler Kullanilarak Yüz Önleştirme


Saribaş H., Turgut K., Cengiz M., Aydemir H., Yeşildag G., Çevikalp H.

34th Signal Processing and Communications Applications Conference, SIU 2026, İstanbul, Türkiye, 7 - 10 Temmuz 2026, (Tam Metin Bildiri)

  • Yayın Türü: Bildiri / Tam Metin Bildiri
  • Doi Numarası: 10.1109/siu71813.2026.11636507
  • Basıldığı Şehir: İstanbul
  • Basıldığı Ülke: Türkiye
  • Anahtar Kelimeler: Face Frontalization, Self-attention Mechanism, Transformer Architecture
  • Eskişehir Osmangazi Üniversitesi Adresli: Evet

Özet

This study proposes a novel Transformer-based framework for synthesizing high-quality frontal face images from profile views. Unlike conventional methods, the proposed approach FaceFrontViT leverages the self-attention mechanism to adaptively focus on sparse facial regions - such as the eyes, nose, and mouth - effectively modeling pose variations and long- range pixel dependencies based on structural symmetry. Within the framework, input images are transformed into patch-based representations incorporating linear embedding and positional encoding, which are subsequently processed through a Transformer encoder architecture. The encoder outputs are further refined by a convolutional reconstruction module to generate the final frontal images. Experimental results on Honda/UCSD recognition datasets demonstrate that our methodology outperforms existing GAN-based approaches and consistently produces high-fidelity frontal results.