TY - GEN
T1 - Automatic Speech Recognition Model Adaptation to Medical Domain Using Untranscribed Audio
AU - Salimbajevs, Askars
AU - Kapočiūtė-Dzikienė, Jurgita
N1 - Publisher Copyright:
© 2022, The Author(s), under exclusive license to Springer Nature Switzerland AG.
PY - 2022
Y1 - 2022
N2 - Automatic speech recognition (ASR) technologies can provide significant efficiency gains in the health sector, by saving time and financial resources, allowing specialists to shift more time to high-value activities. Creating customized ASR models requires domain- and task-related transcribed speech data. Unfortunately, producing such data usually is too expensive for medical institutions: it requires a lot of financial, human resources, and expertise. Consequently, his paper explores a semi-supervised medical domain adaptation method for the Latvian language that benefits from the untranscribed speech recordings. For the initial model, we use the currently available general-purpose hybrid ASR system with the core of a lattice-free maximum mutual information method used to train its acoustic model. The initial system is applied to the domain-related untranscribed data to extract sequences of pseudo-labels. Such automatic transcriptions are later added to the supervised and used together to update the acoustic model. To improve our ASR system further, we have also updated its language model with additional in-domain texts. We have achieved significant improvements in the quality of speech recognition on all evaluation datasets. On the epicrises, psychiatry, and radiology datasets word error rate (WER) decreased by 39%, 27%–29%, and 21%, respectively.
AB - Automatic speech recognition (ASR) technologies can provide significant efficiency gains in the health sector, by saving time and financial resources, allowing specialists to shift more time to high-value activities. Creating customized ASR models requires domain- and task-related transcribed speech data. Unfortunately, producing such data usually is too expensive for medical institutions: it requires a lot of financial, human resources, and expertise. Consequently, his paper explores a semi-supervised medical domain adaptation method for the Latvian language that benefits from the untranscribed speech recordings. For the initial model, we use the currently available general-purpose hybrid ASR system with the core of a lattice-free maximum mutual information method used to train its acoustic model. The initial system is applied to the domain-related untranscribed data to extract sequences of pseudo-labels. Such automatic transcriptions are later added to the supervised and used together to update the acoustic model. To improve our ASR system further, we have also updated its language model with additional in-domain texts. We have achieved significant improvements in the quality of speech recognition on all evaluation datasets. On the epicrises, psychiatry, and radiology datasets word error rate (WER) decreased by 39%, 27%–29%, and 21%, respectively.
KW - Health sector
KW - Hybrid ASR
KW - Latvian language
KW - Medical domain
KW - Semi-supervised
UR - https://www.scopus.com/pages/publications/85134321221
U2 - 10.1007/978-3-031-09850-5_5
DO - 10.1007/978-3-031-09850-5_5
M3 - Conference paper
AN - SCOPUS:85134321221
SN - 9783031098499
T3 - Communications in Computer and Information Science
SP - 65
EP - 79
BT - Digital Business and Intelligent Systems - 15th International Baltic Conference, Baltic DB and IS 2022, Proceedings
A2 - Ivanovic, Mirjana
A2 - Kirikova, Marite
A2 - Niedrite, Laila
PB - Springer Science and Business Media Deutschland GmbH
T2 - 15th International Baltic Conference on Digital Business and Intelligent Systems, Baltic DB and IS 2022
Y2 - 4 July 2022 through 6 July 2022
ER -