Evaluating Deep Learning Architectures for Effective Detection of Manipulated Media across Multiple Datasets

Achraf Ibnouzaher ORCID ,  Noureddine Moumkine ORCID
    Received: 1 February 2026; Revised: 4 March 2026; Accepted: 25 March 2026; Published: 15 July 2026

    Abstract

    The advancement of artificial intelligence has enabled the creation of highly realistic synthetic media raising serious concerns about misinformation, identity manipulation, and digital security. As these technologies continue to evolve and become more accessible the potential misuse of manipulated media increases significantly making it easier to spread deceptive content across digital platforms. This growing threat highlights the urgent need for reliable and efficient deepfake detection techniques capable of identifying manipulated content and limiting its potential societal and technological impact. Deep Learning (DL) has emerged as a promising solution for automatically identifying manipulated multimedia content. In this study, we harnessed Transfer Learning (TL) techniques to fine-tune and evaluate the performance of seven different DL models including EfficientNetB5, XceptionNet, EfficientNet3D, Vision Transformer (ViT), CrossViT, ViViT, and CoAtNet. These models were trained and evaluated on three widely used deepfake detection datasets: FaceForensics++, the DeepFake Detection Challenge (DFDC), and CelebDF. Our experimental results demonstrate that hybrid approaches combining Convolutional Neural Networks (CNNs) with Vision Transformer-based architectures outperform individual models in detecting manipulated media. Furthermore, the findings indicate that models trained on the FaceForensics++ and DFDC datasets exhibit stronger generalization capabilities when applied to previously unseen deepfake samples.

    Keywords

    References

      ×