跳到主要导航 跳到搜索 跳到主要内容

A sequential contrastive learning framework for robust dysarthric speech recognition

  • Lidan Wu
  • , Daoming Zong
  • , Shiliang Sun
  • , Jing Zhao*
  • *此作品的通讯作者
  • East China Normal University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Dysarthria is a manifestation of disruption in the neuromuscular physiology resulting in uneven, slow, slurred, harsh, or quiet speech. Despite the remarkable progress of automatic speech recognition (ASR), it poses great challenges in developing stable ASR for dysarthric individuals due to the high intra- and inter-speaker variations and data deficiency. In this paper, we propose a contrastive learning framework for robust dysarthric speech recognition (DSR) by capturing the dysarthric speech variability. Several speech data augmentation strategies are explored to form two branches of the framework, meanwhile alleviating the scarcity of dysarthria data. We also develop an efficient projection head acting on a sequence of learned hidden representations for defining contrastive loss. Experiment results on DSR demonstrate that the model is better than or comparable to the supervised baseline.

源语言英语
主期刊名2021 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2021 - Proceedings
出版商Institute of Electrical and Electronics Engineers Inc.
7303-7307
页数5
ISBN(电子版)9781728176055
DOI
出版状态已出版 - 2021
活动2021 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2021 - Virtual, Toronto, 加拿大
期限: 6 6月 202111 6月 2021

丛书

姓名ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
2021-June
ISSN(印刷版)1520-6149

会议

会议2021 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2021
国家/地区加拿大
Virtual, Toronto
时期6/06/2111/06/21

学术指纹

探究 'A sequential contrastive learning framework for robust dysarthric speech recognition' 的科研主题。它们共同构成独一无二的学术指纹。

引用此