Classifying Synthetic Facial Images:
A Quantitative and Qualitative Approach with AI
DOI:
https://doi.org/10.22456/2175-2745.150810Keywords:
Xception, Grad-CAM++, synthetic faces, image classifyingAbstract
Through artificial intelligence (AI) techniques, synthetic content is generated, which is an increasingly critical issue for digital security. This paper aims to classify real and synthetic facial images using a state-of-the-art AI model to analyze, both quantitatively and qualitatively, its performance and the most discriminating features across the various data transformations used for this purpose in current literature. In this scenario, the study utilized three public benchmark datasets: two for generating fake images (Stable Diffusion and StyleGAN2) and one for real images (FFHQ), employing the Xception architecture with the Grad-CAM++ viewer. Overall, the results indicate that the central region of the image is the most discriminative one, whereas random cropping produced the worst performance. Future work aims to explore other explainable Artificial Intelligence (XAI) methods, as well as other training and validation protocols, such as cross-testing between synthetic databases.
Downloads
References
[1] GUPTA, P. et al. Generative AI: A systematic review using topic modelling techniques. Data and Information Management, Elsevier, p. 100066, 2024.
[2] KAUSHAL, A.; KUMAR, S.; KUMAR, R. A review on deepfake generation and detection: bibliometric analysis. Multimedia Tools and Applications, Springer, p. 1–41, 2024.
[3] GOODFELLOW, I. et al. Generative adversarial nets. In: GHAHRAMANI, Z. et al. (Ed.). Advances in Neural Information Processing Systems. Curran Associates, Inc., 2014. v. 27. Disponível em: https://proceedings.neurips.cc/paper_files/paper/2014/file/5ca3e9b122f61f8f06494c97b1afccf3-Paper.pdf.
[4] MITRA, A. et al. A machine learning based approach for deepfake detection in social media through key video frame extraction. SN Computer Science, Springer, v. 2, n. 2, p. 1–18, 2021.
[5] TOLOSANA, R. et al. Deepfakes detection across generations: Analysis of facial regions, fusion, and performance evaluation. Engineering Applications of Artificial Intelligence, Elsevier, v. 110, p. 104673, 2022.
[6] PAPA, L. et al. On the use of stable diffusion for creating realistic faces: from generation to detection. In: IEEE. 2023 11th International Workshop on Biometrics and Forensics (IWBF). [S.l.], 2023. p. 1–6.
[7] ROSSLER, A. et al. Faceforensics++: Learning to detect manipulated facial images. In: Proceedings of the IEEE/CVF international conference on computer vision. [S.l.: s.n.], 2019. p. 1–11.
[8] LI, Y. et al. Celeb-df: A large-scale challenging dataset for deepfake forensics. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. [S.l.: s.n.], 2020. p. 3207–3216.
[9] DOLHANSKY, B. et al. The deepfake detection challenge (dfdc) preview dataset. arXiv preprint arXiv:1910.08854, 2019.
[10] LI, Y.; CHANG, M.-C.; LYU, S. In ictu oculi: Exposing AI created fake videos by detecting eye blinking. In: IEEE. 2018 IEEE international workshop on information forensics and security (WIFS). [S.l.], 2018. p. 1–7.
[11] CHOLLET, F. Xception: Deep learning with depthwise separable convolutions. In: Proceedings of the IEEE conference on computer vision and pattern recognition. [S.l.: s.n.], 2017. p. 1251–1258.
[12] ROMBACH, R. et al. High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). [S.l.: s.n.], 2022. p. 10684–10695.
[13] KARRAS, T.; LAINE, S.; AILA, T. A style-based generator architecture for generative adversarial networks. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. [S.l.: s.n.], 2019. p. 4401–4410.
[14] KARRAS, T. et al. Analyzing and improving the image quality of StyleGAN. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. [S.l.: s.n.], 2020. p. 8110–8119.
[15] KARRAS, T. et al. Alias-free generative adversarial networks. In: Proc. NeurIPS. [S.l.: s.n.], 2021.
[16] SIMONYAN, K.; ZISSERMAN, A. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
[17] HE, K. et al. Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition. [S.l.: s.n.], 2016. p. 770–778.
[18] HOWARD, A. G. et al. MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017.
[19] DOSOVITSKIY, A. et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
[20] CHATTOPADHAY, A. et al. Grad-CAM++: Generalized gradient-based visual explanations for deep convolutional networks. In: IEEE. 2018 IEEE winter conference on applications of computer vision (WACV). [S.l.], 2018. p. 839–847.
[21] SCHUHMANN, C. et al. LAION-400M: Open dataset of CLIP-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021.
[22] KARRAS, T. et al. Progressive growing of GANs for improved quality, stability, and variation. arXiv preprint arXiv:1710.10196, 2018.
[23] YU, F. et al. LSUN: Construction of a large-scale image dataset using deep learning with humans in the loop. arXiv preprint arXiv:1506.03365, 2015.
[24] BALTRUŠAITIS, T.; ROBINSON, P.; MORENCY, L.-P. OpenFace: an open source facial behavior analysis toolkit. In: IEEE. 2016 IEEE winter conference on applications of computer vision (WACV). [S.l.], 2016. p. 1–10.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Fernanda Goyo Tamanaka, Lucas Fontes Buzuti, Carlos Eduardo Thomaz

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Autorizo aos editores a publicação de meu artigo, caso seja aceito, em meio eletrônico de acordo com as regras do Public Knowledge Project.













