| [1] |
国家互联网信息办公室. 互联网信息服务深度合成管理规定[EB/OL]. (2022-12-11)[2026-04-23]. https://www.gov.cn/zhengce/zhengceku/2022-12/12/content_5731431.htm.
|
| [2] |
奇安信集团. 2024人工智能安全报告[R/OL]. https://www.qianxin.com/北京:奇安信集团, 2024.
|
| [3] |
新华社. “三只羊”公司使用AI伪造视频音频事件调查报道[N/OL]. 新华社, 2024. https://www.xinhuanet.com.
|
| [4] |
腾讯新闻. 特朗普幕僚长深陷AI语音诈骗案: 黑客窃取其通讯录, 冒充身份接触政要并要求商人打钱[N/OL]. 腾讯新闻. [2025-05-30]. https://news.qq.com/rain/a/20250530A08QZH00.
|
| [5] |
许裕雄, 李斌, 谭舜泉, 等. 语音深度伪造及其检测技术研究进展[J]. 中国图象图形学报, 2024, 29(8): 2236-2268.
|
| [6] |
JI Z, LI Z Y, LI P, et al. Ensemble learning for countermeasure of audio replay spoofing attack in ASVspoof 2017[C]// Interspeech 2017: Annual Conference of the International Speech Communication Association. Stockholm, Sweden: ISCA, 2017: 87-91.
|
| [7] |
PHAM L, LAM P, TRAN D, et al. A comprehensive survey with critical analysis for deepfake speech detection[J]. Computer Science Review, 2025, 57: 100757.
doi: 10.1016/j.cosrev.2025.100757
|
| [8] |
KUMAR S. Real-time implementation and performance evaluation of speech classifiers in speech analysis-synthesis[J]. ETRI Journal, 2021, 43(1): 82-94.
doi: 10.4218/etrij.2019-0364
|
| [9] |
WU Z, KINNUNEN T, EVANS N, et al. ASVspoof 2015: Automatic speaker verification spoofing and countermeasures challenge evaluation plan[J]. Training, 2014, 10(15): 3750.
|
| [10] |
WANG X, YAMAGISHI J, TODISCO M, et al. ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech[J]. Computer Speech & Language, 2020, 64: 101114.
|
| [11] |
LIU X, WANG X, SAHIDULLAH M, et al. Asvspoof 2021: Towards spoofed and deepfake speech detection in the wild[J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2023, 31: 2507-2522.
doi: 10.1109/TASLP.2023.3285283
|
| [12] |
DELGADO H, EVANS N, JUNG J, et al. ASVspoof 5 Evaluation Plan[EB/OL]. (2024-06-28)[2026-04-23]. https://www.asvspoof.org/file/ASVspoof5_Evaluation_Plan_Phase2.pdf.
|
| [13] |
YI J, FU R, TAO J, et al. Add 2022: the first audio deep synthesis detection challenge[C]// ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022: 9216-9220.
|
| [14] |
YI J, TAO J, FU R, et al. ADD 2023: The Second Audio Deepfake Detection Challenge[J]. arXiv preprint arXiv:2305.13774, 2023.
|
| [15] |
FRANK J, SCHÖNHERR L. Wavefake: A data set to facilitate audio deepfake detection[J]. arXiv preprint arXiv:2111.02813, 2021.
|
| [16] |
REIMAO R, TZERPOS V. FoR: A dataset for synthetic speech detection[C]// 2019 International Conference on Speech Technology and Human-Computer Dialogue (SpeD). IEEE, 2019: 1-10.
|
| [17] |
SALVI D, HOSLER B, BESTAGINI P, et al. TIMIT-TTS: A text-to-speech dataset for multimodal synthetic media detection[J]. IEEE Access, 2023, 11: 50851-50866
doi: 10.1109/ACCESS.2023.3276480
|
| [18] |
YAN Z, ZHAO Y, WANG H. VoiceWukong: Benchmarking deepfake voice detection[J]. arXiv preprint arXiv:2409.06348, 2024.
|
| [19] |
MÜLLER N M, CZEMPIN P, DIECKMANN F, et al. Does audio deepfake detection generalize?[J]. arXiv preprint arXiv:2203.16263, 2022.
|
| [20] |
DOLHANSKY B, BITTON J, PFLAUM B, et al. The deepfake detection challenge (DFDC) dataset[J]. arXiv preprint arXiv:2006.07397, 2020.
|
| [21] |
KWON P, YOU J, NAM G, et al. KoDF: A large-scale korean deepfake detection dataset[C]// Proceedings of the IEEE/CVF international conference on computer vision. 2021: 10744-10753.
|
| [22] |
KHALID H, TARIQ S, KIM M, et al. FakeAVCeleb: A novel audio-video multimodal deepfake dataset[J]. arXiv preprint arXiv:2108.05080, 2021.
|
| [23] |
CAI Z, STEFANOV K, DHALL A, et al. Do you really mean that?content driven audio-visual deepfake dataset and multimodal method for temporal forgery localization[C]// 2022 International Conference on Digital Image Computing: Techniques and Applications (DICTA). IEEE, 2022: 1-10.
|
| [24] |
CAI Z, GHOSH S, ADATIA A P, et al. AV-Deepfake1M:A large-scale LLM-driven audio-visual deepfake dataset[C]// Proceedings of the 32nd ACM International Conference on Multimedia. 2024: 7414-7423.
|
| [25] |
JUNG J, WU Y, WANG X, et al. SpoofCeleb: Speech deepfake detection and SASV in the wild[J]. IEEE Open Journal of Signal Processing, 2025, 6: 68-77.
doi: 10.1109/OJSP.2025.3529377
|
| [26] |
HUANG W, GU Y, WANG Z, et al. SpeechFake: A large-scale multilingual speech deepfake dataset toward cutting-edge speech generation methods[C]// ACL 2025: 63rd Annual Meeting of the Association for Computational Linguistics. Vienna, Austria: Association for Computational Linguistics, 2025: 9985-9998.
|
| [27] |
LI X F, LI K, ZHENG Y F, et al. Safeear: Content privacy-preserving audio deepfake detection[C]// CCS 2024: ACM SIGSAC Conference on Computer and Communications Security. Salt Lake City, UT, USA: ACM, 2024: 3585-3599.
|
| [28] |
MA H, YI J, WANG C, et al. CFAD: A Chinese dataset for fake audio detection[J]. Speech Communication, 2024, 164: 103122.
doi: 10.1016/j.specom.2024.103122
|
| [29] |
HOANG V, PHAM V T, XUAN H N, et al. VSASV: A Vietnamese dataset for spoofing-aware speaker verification[C]// Interspeech 2024:Annual Conference of the International Speech Communication Association. Kos, Greece: ISCA, 2024: 4288-4292.
|
| [30] |
LAVRENTYEVA G, NOVOSELOV S, Tseren A, et al. STC antispoofing systems for the ASVspoof2019 challenge[J]. arXiv preprint arXiv:1904.05576, 2019.
|
| [31] |
LEI Z, WEN Y, YANG Y, et al. Group GMM-ResNet for Detection of Synthetic Speech Attacks[C]// Proc. Interspeech 2023. 2023: 3187-3191.
|
| [32] |
LEI Z, YAN H, LIU C, et al. GMM-ResNet2:Ensemble of Group ResNet Networks for Synthetic Speech Detection[C]// ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2024: 12101-12105.
|
| [33] |
MALTBY H, WALL J, GLACKIN C, et al. A frequency bin analysis of distinctive ranges between human and deepfake generated voices[C]// IJCNN 2024: International Joint Conference on Neural Networks. Yokohama, Japan: IEEE, 2024: 1-7.
|
| [34] |
TAK H, JUNG J, PATINO J, et al. End-to-end spectro-temporal graph attention networks for speaker verification anti-spoofing and speech deepfake detection[J]. arXiv preprint arXiv:2107.12710, 2021.
|
| [35] |
DESPLANQUES B, THIENPONDT J, DEMUYNCK K. ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification[J]. arXiv preprint arXiv:2005.07143, 2020.
|
| [36] |
TAK H, PATINO J, TODISCO M, et al. End-to-end anti-spoofing with RawNet2[C]// ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021: 6369-6373.
|
| [37] |
LIU J, SU Z, HUANG H, et al. TranssionADD: A multi-frame reinforcement based sequence tagging model for audio deepfake detection[J]. arXiv preprint arXiv:2306.15212, 2023.
|
| [38] |
JUNG J W, HEO H S, TAK H, et al. AASIST: Audio anti-spoofing using integrated spectro-temporal graph attention networks[C]//ICASSP 2022: IEEE International Conference on Acoustics, Speech and Signal Processing.Singapore:IEEE, 2022: 6367-6371.
|
| [39] |
ZHANG Q, WEN S, HU T. Audio deepfake detection with self-supervised XLS-R and SLS classifier[C]// Proceedings of the 32nd ACM International Conference on Multimedia. 2024: 6765-6773.
|
| [40] |
TRAN H M, LOLIVE D, SINI A, et al. Multi-level SSL Feature Gating for Audio Deepfake Detection[C]// Proceedings of the 33rd ACM International Conference on Multimedia. 2025: 11766-11775.
|
| [41] |
TRUONG D T, TAO R, NGUYEN T, et al. Temporal-channel modeling in multi-head self-attention for synthetic speech detection[J]. arXiv preprint arXiv:2406. 17376, 2024.
|
| [42] |
LI L, LU T, MA X, et al. Voice deepfake detection using the self-supervised pre-training model hubert[J]. Applied Sciences, 2023, 13(14): 8488.
doi: 10.3390/app13148488
|
| [43] |
XIAO Y, DAS R K. XLSR-Mamba: A dual-column bidirectional state space model for spoofing attack detection[J]. IEEE Signal Processing Letters, 2025, 32: 1276-1280.
doi: 10.1109/LSP.2025.3547861
|
| [44] |
ZHANG Y, LU J, SHANG Z, et al. Improving short utterance anti-spoofing with aasist2[C]// ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024: 11636-11640.
|
| [45] |
GUO Y, HUANG H, CHEN X, et al. Audio deepfake detection with self-supervised wavlm and multi-fusion attentive classifier[C]// ICASSP 2024-2024 IEEE International Conference on Acoustics,Speech and Signal Processing(ICASSP).IEEE, 2024: 12702-12706.
|
| [46] |
DOWERAH S, KULKARNI A, KULKARNI A, et al. Speech df arena: A leaderboard for speech deepfake detection models[J]. arXiv preprint arXiv:2509. 02859, 2025.
|
| [47] |
WANG X, YAMAGISHI J. A comparative study on recent neural spoofing countermeasures for synthetic speech detection[J]. arXiv preprint arXiv:2103.11326, 2021.
|
| [48] |
ZHANG Y, JIANG F, DUAN Z. One-class learning towards synthetic voice spoofing detection[J]. IEEE Signal Processing Letters, 2021, 28: 937-941.
doi: 10.1109/LSP.2021.3076358
|
| [49] |
REDDY T U K, VARUN S C, SREEKANTH K P K S, et al. Evince the artifacts of spoof speech by blending vocal tract and voice source features[J]. arXiv preprint arXiv:2212.02013, 2022.
|
| [50] |
CUCCOVILLO L, GERHARDT M, AICHROTH P. Audio spectrogram transformer for synthetic speech detection via speech formant analysis[C]// WIFS 2023: IEEE International Workshop on Information Forensics and Security. Nürnberg, Germany: IEEE, 2023: 1-6.
|
| [51] |
SHIN H, HEO J, KIM J, et al. Hm-conformer: A conformer-based audio deepfake detection system with hierarchical pooling and multi-level classification token aggregation methods[C]// ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024: 10581-10585.
|
| [52] |
TAK H, KAMBLE M, PATINO J, et al. RawBoost: A raw data boosting and augmentation method applied to automatic speaker verification anti-spoofing[C]// ICASSP 2022: IEEE International Conference on Acoustics, Speech and Signal Processing. Singapore: IEEE, 2022: 6382-6386.
|
| [53] |
PAN Z, LIU T, SAILOR H B, et al. Attentive merging of hidden embeddings from pre-trained speech model for anti-spoofing detection[J]. arXiv preprint arXiv:2406.10283, 2024.
|