Speaker recognition based on deep learning: An overview. 2021

Zhongxin Bai, and Xiao-Lei Zhang
Center of Intelligent Acoustics and Immersive Communications (CIAIC) and the School of Marine Science and Technology, Northwestern Polytechnical University, Xi'an Shaanxi 710072, China. Electronic address: zxbai@mail.nwpu.edu.cn.

Speaker recognition is a task of identifying persons from their voices. Recently, deep learning has dramatically revolutionized speaker recognition. However, there is lack of comprehensive reviews on the exciting progress. In this paper, we review several major subtasks of speaker recognition, including speaker verification, identification, diarization, and robust speaker recognition, with a focus on deep-learning-based methods. Because the major advantage of deep learning over conventional methods is its representation ability, which is able to produce highly abstract embedding features from utterances, we first pay close attention to deep-learning-based speaker feature extraction, including the inputs, network structures, temporal pooling strategies, and objective functions respectively, which are the fundamental components of many speaker recognition subtasks. Then, we make an overview of speaker diarization, with an emphasis of recent supervised, end-to-end, and online diarization. Finally, we survey robust speaker recognition from the perspectives of domain adaptation and speech enhancement, which are two major approaches of dealing with domain mismatch and noise problems. Popular and recently released corpora are listed at the end of the paper.

UI MeSH Term Description Entries
D009622 Noise Any sound which is unwanted or interferes with HEARING other sounds. Noise Pollution,Noises,Pollution, Noise
D000077321 Deep Learning Supervised or unsupervised machine learning methods that use multiple layers of data representations generated by nonlinear transformations, instead of individual task-specific ALGORITHMS, to build and train neural network models. Hierarchical Learning,Learning, Deep,Learning, Hierarchical
D049250 Speech Recognition Software Software capable of recognizing dictation and transcribing the spoken words into written text. Voice Recognition Software,Recognition Software, Speech,Recognition Software, Voice,Software, Speech Recognition,Software, Voice Recognition

Related Publications

Zhongxin Bai, and Xiao-Lei Zhang
October 2018, IEEE/ACM transactions on audio, speech, and language processing,
Zhongxin Bai, and Xiao-Lei Zhang
February 2023, Sensors (Basel, Switzerland),
Zhongxin Bai, and Xiao-Lei Zhang
March 2022, Computer methods and programs in biomedicine,
Zhongxin Bai, and Xiao-Lei Zhang
June 2003, Hang tian yi xue yu yi xue gong cheng = Space medicine & medical engineering,
Zhongxin Bai, and Xiao-Lei Zhang
November 2020, Neural networks : the official journal of the International Neural Network Society,
Zhongxin Bai, and Xiao-Lei Zhang
June 2023, Sensors (Basel, Switzerland),
Zhongxin Bai, and Xiao-Lei Zhang
January 2022, Computational intelligence and neuroscience,
Zhongxin Bai, and Xiao-Lei Zhang
February 2024, Mathematical biosciences and engineering : MBE,
Copied contents to your clipboard!