跳到主要导航 跳到搜索 跳到主要内容

Speaker Extraction with Detection of Presence and Absence of Target Speakers

  • Ke Zhang
  • , Marvin Borsdorf
  • , Zexu Pan
  • , Haizhou Li
  • , Yangjie Wei
  • , Yi Wang
  • Northeastern University China
  • National University of Singapore
  • University of Bremen
  • The Chinese University of Hong Kong, Shenzhen

科研成果: 期刊稿件会议文章同行评审

摘要

Target speaker extraction extracts a target voice from a given cocktail party mixture signal. Most studies are restricted to conditions in which the target speaker is present in the mixture (PT), which often fail when the target speaker is absent (AT). Training on both PT and AT situations helps, but degrades the PT performance as the model intrinsically tries to detect the target presence. We propose a new model, called TSEJoint, that jointly performs target speaker detection and extraction. Both tasks share the low-level modules, allowing the detection branch to use a pre-separated signal and keeping the overall processing pipeline length similar, while at the high-level they have different branches to ensure the performance of each task. We evaluate our proposed methods under PT and AT conditions comprising one and two talkers. The TSEJoint model shows better extraction performance under the PT condition and better detection performance on all conditions compared with the baseline.

源语言英语
页(从-至)3714-3718
页数5
期刊Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
2023-August
DOI
出版状态已出版 - 2023
已对外发布
活动24th Annual conference of the International Speech Communication Association, Interspeech 2023 - Dublin, 爱尔兰
期限: 20 8月 202324 8月 2023

学术指纹

探究 'Speaker Extraction with Detection of Presence and Absence of Target Speakers' 的科研主题。它们共同构成独一无二的学术指纹。

引用此