数据与计算发展前沿 ›› 2026, Vol. 8 ›› Issue (4): 111-122.

CSTR: 32002.14.jfdc.CN10-1649/TP.2026.04.008

doi: 10.11871/jfdc.issn.2096-742X.2026.04.008

• • 上一篇    下一篇

语音深度伪造检测:技术演进与展望

王丹琳1,2(),唐云祁2,*()   

  1. 1 北京警察学院刑事科学技术系北京 102202
    2 中国人民公安大学侦查学院北京 100038
  • 收稿日期:2026-03-04 出版日期:2026-08-20 发布日期:2026-08-21
  • 通讯作者: 唐云祁(E-mail: tangyunqi@ppsuc.edu.cn
  • 作者简介:王丹琳,北京警察学院刑事科学技术系,讲师,主要研究方向为视听检验技术、智能图像识别与理解。
    本文承担工作:负责系统性文献调研、论文的撰写。
    WANG Danlin is a lecturer in the Department of Criminal Science and Technology, Beijing Police College. Her research interests include forensic audio-visual examination technology, intelligent image recognition and interpretation.
    In this paper, she is responsible for systematic literature review and manuscript writing.
    E-mail: wangdl.1101@163.com|唐云祁,中国人民公安大学侦查学院,教授,博士生导师,主要研究方向为电子数据检验、智能图像识别与理解。
    本文承担工作为:负责论文框架设计,论文修改与审定。
    TANG Yunqi is a professor and doctoral supervisor at the School of Criminal Investigation, People’s Public Security University of China. His research interests include electronic data examination and forensics, intelligent image recognition and interpretation.
    In this paper, he is mainly responsible for designing the manuscript framework, revising the manuscript, and giving final approval.
    E-mail: tangyunqi@ppsuc.edu.cn
  • 基金资助:
    北京市公安局科研项目(2026CX27012)

Deepfake Speech Detection: Technological Evolution and Challenges

WANG Danlin1,2(),TANG Yunqi2,*()   

  1. 1 Department of Criminal Science and Technology, Beijing Police College, Beijing 102202, China
    2 School of Criminal Investigation, People’s Public Security University of China, Beijing 100038, China
  • Received:2026-03-04 Online:2026-08-20 Published:2026-08-21

摘要:

【目的】 系统梳理语音深度伪造检测技术的演进脉络与关键问题,分析其发展趋势及技术挑战。【文献范围】 基于2021—2025年国内外约180篇相关研究文献,涵盖主流公开数据集与代表性检测技术。【方法】 首先总结现有语音深度伪造检测数据集的构建特点与局限;随后从空间特征建模、时序动态建模、多机制融合与协同建模、自监督预训练驱动建模四个阶段归纳检测技术的演进,并对代表性模型的性能表现进行比较分析。【结果】 研究表明,语音深度伪造检测技术正由单一建模向多机制协同与预训练模型融合发展,检测性能持续提升,但在跨域泛化、真实环境鲁棒性、可解释性及高效部署等方面仍面临挑战。【结论】 未来研究应加强高真实性数据集建设,完善统一评测体系,并探索具有良好跨域泛化能力和可解释性的检测模型,以提升该技术在复杂真实场景中的鲁棒性与应用能力。

关键词: 语音深度伪造检测, 数据集, 空间特征建模, 时序动态建模, 多机制融合, 自监督预训练模型

Abstract:

[Purpose] This paper systematically reviews the evolutionary trajectory and key issues of speech deepfake detection technologies, and analyzes their development trends and technical challenges. [Literature Scope] The review is based on approximately 180 related studies published between 2021 and 2025, covering mainstream public datasets and representative detection technologies. [Methods] First, the construction characteristics and limitations of existing speech deepfake detection datasets are summarized. Then, the evolution of detection technologies is analyzed across four stages, including spatial feature modeling, temporal dynamic modeling, multi-mechanism fusion and collaborative modeling, and self-supervised pretraining-driven modeling. Finally, the performance of representative models is comparatively evaluated. [Results] The review indicate that speech deepfake detection technologies are evolving from single-mechanism modeling toward multi-mechanism collaboration and integration with pre-trained models, leading to continuous improvements in detection performance. However, challenges remain in cross-domain generalization, robustness in real-world environments, interpretability, and efficient deployment. [Conclusions] Future research should focus on constructing high-fidelity datasets, establishing unified evaluation protocols, and developing detection models with strong cross-domain generalization and interpretability, thereby improving the robustness and practical applicability of these technologies in complex real-world scenarios.

Key words: speech deepfake detection, datasets, spatial feature modeling, temporal dynamic modeling, multi-mechanism fusion, self-supervised pretraining