A Spatial and Temporal Hybrid Deep Learning Framework for Deep Fake Detection
DOI:
https://doi.org/10.47392/IRJAEH.2026.0281Keywords:
Deep Fake Detection, Digital Media Security Hybrid Deep Learning, Spatial Feature Extraction, Temporal Feature Analysis, Multi-Modal Media Authentication.Abstract
The rapid advancement of artificial intelligence has significantly improved multimedia generation technologies, but it has also led to the emergence of deep fake content that threatens digital media authenticity and security. This paper proposes a Hybrid Deep Learning Framework for Deep Fake Detection that integrates spatial and temporal feature extraction techniques to identify manipulated images, videos, and audio content. For image-based detection, advanced object detection models including YOLOv8, YOLOv10, Fast-RCNN, and EfficientDet are employed to analyze facial inconsistencies and spatial artifacts. For video deep fake detection, InceptionNet is integrated with a Gated Recurrent Unit (GRU) network to capture both frame-level spatial features and sequential temporal dependencies. Additionally, a Convolutional Neural Network (CNN) model is utilized for detecting synthetic audio manipulations through sound pattern analysis. The system is deployed through a Flask- based web interface that allows users to upload multimedia files and receive authenticity predictions. Performance evaluation is conducted using qualitative and quantitative metrics such as accuracy, precision, recall, F1-score, and confusion matrix analysis. Expected results demonstrate high detection accuracy and robustness across different media types. The proposed framework contributes to digital media security by offering a scalable and practical solution for automated deep fake identification.
Downloads
Downloads
Published
Issue
Section
License
Copyright (c) 2026 International Research Journal on Advanced Engineering Hub (IRJAEH)

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
.