TY - JOUR
T1 - Speaker Extraction with Detection of Presence and Absence of Target Speakers
AU - Zhang, Ke
AU - Borsdorf, Marvin
AU - Pan, Zexu
AU - Li, Haizhou
AU - Wei, Yangjie
AU - Wang, Yi
N1 - Publisher Copyright:
© 2023 International Speech Communication Association. All rights reserved.
PY - 2023
Y1 - 2023
N2 - Target speaker extraction extracts a target voice from a given cocktail party mixture signal. Most studies are restricted to conditions in which the target speaker is present in the mixture (PT), which often fail when the target speaker is absent (AT). Training on both PT and AT situations helps, but degrades the PT performance as the model intrinsically tries to detect the target presence. We propose a new model, called TSEJoint, that jointly performs target speaker detection and extraction. Both tasks share the low-level modules, allowing the detection branch to use a pre-separated signal and keeping the overall processing pipeline length similar, while at the high-level they have different branches to ensure the performance of each task. We evaluate our proposed methods under PT and AT conditions comprising one and two talkers. The TSEJoint model shows better extraction performance under the PT condition and better detection performance on all conditions compared with the baseline.
AB - Target speaker extraction extracts a target voice from a given cocktail party mixture signal. Most studies are restricted to conditions in which the target speaker is present in the mixture (PT), which often fail when the target speaker is absent (AT). Training on both PT and AT situations helps, but degrades the PT performance as the model intrinsically tries to detect the target presence. We propose a new model, called TSEJoint, that jointly performs target speaker detection and extraction. Both tasks share the low-level modules, allowing the detection branch to use a pre-separated signal and keeping the overall processing pipeline length similar, while at the high-level they have different branches to ensure the performance of each task. We evaluate our proposed methods under PT and AT conditions comprising one and two talkers. The TSEJoint model shows better extraction performance under the PT condition and better detection performance on all conditions compared with the baseline.
KW - absent speaker
KW - cocktail party problem
KW - selective auditory attention
KW - speaker detection
KW - target speaker extraction
UR - https://www.scopus.com/pages/publications/85171522920
U2 - 10.21437/Interspeech.2023-655
DO - 10.21437/Interspeech.2023-655
M3 - 会议文章
AN - SCOPUS:85171522920
SN - 2308-457X
VL - 2023-August
SP - 3714
EP - 3718
JO - Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
JF - Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
T2 - 24th Annual conference of the International Speech Communication Association, Interspeech 2023
Y2 - 20 August 2023 through 24 August 2023
ER -