An Audio-Visual Attention Based Multimodal Network for Fake Talking Face Videos Detection

by   Ganglai Wang, et al.

DeepFake based digital facial forgery is threatening the public media security, especially when lip manipulation has been used in talking face generation, the difficulty of fake video detection is further improved. By only changing lip shape to match the given speech, the facial features of identity is hard to be discriminated in such fake talking face videos. Together with the lack of attention on audio stream as the prior knowledge, the detection failure of fake talking face generation also becomes inevitable. Inspired by the decision-making mechanism of human multisensory perception system, which enables the auditory information to enhance post-sensory visual evidence for informed decisions output, in this study, a fake talking face detection framework FTFDNet is proposed by incorporating audio and visual representation to achieve more accurate fake talking face videos detection. Furthermore, an audio-visual attention mechanism (AVAM) is proposed to discover more informative features, which can be seamlessly integrated into any audio-visual CNN architectures by modularization. With the additional AVAM, the proposed FTFDNet is able to achieve a better detection performance on the established dataset (FTFDD). The evaluation of the proposed work has shown an excellent performance on the detection of fake talking face videos, which is able to arrive at a detection rate above 97


page 1

page 2

page 6

page 8


FTFDNet: Learning to Detect Talking Face Video Manipulation with Tri-Modality Interaction

DeepFake based digital facial forgery is threatening public media securi...

On the Detection of Digital Face Manipulation

Detecting manipulated facial images and videos is an increasingly import...

Detection of Cross-Dataset Fake Audio Based on Prosodic and Pronunciation Features

Existing fake audio detection systems perform well in in-domain testing,...

Attention-Based Lip Audio-Visual Synthesis for Talking Face Generation in the Wild

Talking face generation with great practical significance has attracted ...

DeepFakes Evolution: Analysis of Facial Regions and Fake Detection Performance

Media forensics has attracted a lot of attention in the last years in pa...

What's wrong with this video? Comparing Explainers for Deepfake Detection

Deepfakes are computer manipulated videos where the face of an individua...

Exposing Deepfake with Pixel-wise AR and PPG Correlation from Faint Signals

Deepfake poses a serious threat to the reliability of judicial evidence ...

Please sign up or login with your details

Forgot password? Click here to reset