Frontiers of Data and Computing ›› 2026, Vol. 8 ›› Issue (4): 111-122.

CSTR: 32002.14.jfdc.CN10-1649/TP.2026.04.008

doi: 10.11871/jfdc.issn.2096-742X.2026.04.008

Previous Articles     Next Articles

Deepfake Speech Detection: Technological Evolution and Challenges

WANG Danlin1,2(),TANG Yunqi2,*()   

  1. 1 Department of Criminal Science and Technology, Beijing Police College, Beijing 102202, China
    2 School of Criminal Investigation, People’s Public Security University of China, Beijing 100038, China
  • Received:2026-03-04 Online:2026-08-20 Published:2026-08-21

Abstract:

[Purpose] This paper systematically reviews the evolutionary trajectory and key issues of speech deepfake detection technologies, and analyzes their development trends and technical challenges. [Literature Scope] The review is based on approximately 180 related studies published between 2021 and 2025, covering mainstream public datasets and representative detection technologies. [Methods] First, the construction characteristics and limitations of existing speech deepfake detection datasets are summarized. Then, the evolution of detection technologies is analyzed across four stages, including spatial feature modeling, temporal dynamic modeling, multi-mechanism fusion and collaborative modeling, and self-supervised pretraining-driven modeling. Finally, the performance of representative models is comparatively evaluated. [Results] The review indicate that speech deepfake detection technologies are evolving from single-mechanism modeling toward multi-mechanism collaboration and integration with pre-trained models, leading to continuous improvements in detection performance. However, challenges remain in cross-domain generalization, robustness in real-world environments, interpretability, and efficient deployment. [Conclusions] Future research should focus on constructing high-fidelity datasets, establishing unified evaluation protocols, and developing detection models with strong cross-domain generalization and interpretability, thereby improving the robustness and practical applicability of these technologies in complex real-world scenarios.

Key words: speech deepfake detection, datasets, spatial feature modeling, temporal dynamic modeling, multi-mechanism fusion, self-supervised pretraining