Predicting the Reliability of Automatic Speech Recognition in Long-form Recordings
Galama, Azarias (2026)
Galama, Azarias
2026
Bachelor's Programme in Science and Engineering
Tekniikan ja luonnontieteiden tiedekunta - Faculty of Engineering and Natural Sciences
Hyväksymispäivämäärä
2026-04-23
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi:tuni-202604234213
https://urn.fi/URN:NBN:fi:tuni-202604234213
Tiivistelmä
Due to advancements in deep learning and availability of large training data, modern automatic speech recognition (ASR) systems achieve remarkable transcription accuracy under certain benchmarks. However, their performance degrades in case of overlapping speech and when the speech data is collected in noisy naturalistic environments. Therefore, creating a pipeline that predicts the reliability of the outputs of ASR can help mitigate the need for laborious manual inspections, especially in critical domains like child-centered long-form recordings.
This thesis proposes a machine learning pipeline that predicts whether the word error rate (WER) of an utterance is low or high. To achieve this, the continuous WER values were binarized using a threshold of 0.3 into either low-WER or high-WER classes. The proposed classification model was an ensemble model composed of CatBoost, XGBoost, LightGBM, and logistic regression. This model was trained on a feature set comprising ASR-derived features such as aggregate confidence scores from Whisper and WER distance between two Whisper models, and acoustic features such as SNR.
Four child-centered corpora were evaluated in a leave-one-corpus-out strategy. The proposed model on average achieved a precision of 0.83, recall of 0.51, mean WER of 0.15, and median WER of 0.01 in the predicted low-WER class, while the performance across the four corpora was relatively consistent. In addition, the results indicated that mean ASR confidence score and the inter-model transcription difference (WER distance) are especially strong features for predicting transcription reliability. Overall, the proposed approach can help predict reliable ASR outputs and reduce the manual effort required to inspect transcriptions of large-scale naturalistic speech recordings.
This thesis proposes a machine learning pipeline that predicts whether the word error rate (WER) of an utterance is low or high. To achieve this, the continuous WER values were binarized using a threshold of 0.3 into either low-WER or high-WER classes. The proposed classification model was an ensemble model composed of CatBoost, XGBoost, LightGBM, and logistic regression. This model was trained on a feature set comprising ASR-derived features such as aggregate confidence scores from Whisper and WER distance between two Whisper models, and acoustic features such as SNR.
Four child-centered corpora were evaluated in a leave-one-corpus-out strategy. The proposed model on average achieved a precision of 0.83, recall of 0.51, mean WER of 0.15, and median WER of 0.01 in the predicted low-WER class, while the performance across the four corpora was relatively consistent. In addition, the results indicated that mean ASR confidence score and the inter-model transcription difference (WER distance) are especially strong features for predicting transcription reliability. Overall, the proposed approach can help predict reliable ASR outputs and reduce the manual effort required to inspect transcriptions of large-scale naturalistic speech recordings.
Kokoelmat
- Kandidaatintutkielmat [11807]
