Logo

42. Wissenschaftliche Jahrestagung der Deutschen Gesellschaft für Phoniatrie und Pädaudiologie (DGPP)

Deutsche Gesellschaft für Phoniatrie und Pädaudiologie e. V.
16.-19.09.2026
Starnberg

Vortrag

Automated hoarseness severity estimation using high-speed videoendoscopy-synchronized acoustic signals under phonation constraints

P. Patel - University Hospital Erlangen, Division of Phoniatrics and Pediatric Audiology, Department of Otorhinolaryngology Erlangen, Head and Neck Surgery, Erlangen, Germany
J. Donhauser - University Hospital Erlangen, Division of Phoniatrics and Pediatric Audiology, Department of Otorhinolaryngology Erlangen, Head and Neck Surgery, Erlangen, Germany
A. Schützenberger - University Hospital Erlangen, Division of Phoniatrics and Pediatric Audiology, Department of Otorhinolaryngology Erlangen, Head and Neck Surgery, Erlangen, Germany
M. Kunduk - Louisiana State University, Department of Communication Sciences and Disorders, Louisiana, United States
M. Döllinger - University Hospital Erlangen, Division of Phoniatrics and Pediatric Audiology, Department of Otorhinolaryngology Erlangen, Head and Neck Surgery, Erlangen, Germany

Abstract

Background: Hoarseness caused by different voice disorders may exhibit different acoustic signatures, complicating standardized quantitative assessment. Acoustic voice signals recorded during sustained phonation are commonly used for objective assessment of voice quality. Laryngeal high-speed videoendoscopy (HSV) enables simultaneous recording of acoustic signals during phonation. However, the presence of rigid endoscope in the oral cavity results in acoustic signals that may differ from natural voice production. Hence, the goals of this study are to determine an optimal, as short as possible, analysis time interval and analyze the potential of HSV-synchronized acoustic signals for machine learning (ML)-based hoarseness severity estimation.

Materials and methods: Two databases containing normal voices, functional and organic voice disorders were constructed. Database D1 comprises 824 HSV-synchronized acoustic recordings of sustained vowel /i/, while Database D2 includes 804 sustained vowel /a/ recordings from speech therapy sessions. Recording segments of 250 ms, 500 ms, and 1000 ms were analyzed. RBH ratings derived from continuous speech served as ground truth. Subjects were categorized into two hoarseness levels (H < 2 vs. H ≥ 2), while ML-based estimated probability was interpreted as a continuous interval-scaled severity score between 0 and 1. Comprehensive acoustic features were extracted and reduced using ensemble feature selection. ML models of varying complexity (logistic regression, SVM, XGBoost, TabNet) were evaluated across durations.

Results: At 1000 ms, Logistic regression for D1 and XGBoost for D2 achieved the best performance, yielding Spearman rank correlations of 0.62 and 0.75 respectively between predicted severity scores and perceptual ratings. Across models, D2 demonstrated consistently higher performance than D1, with mean differences of approximately 10% in accuracy and 8% in ROC-AUC.

Conclusion: The lower performance of HSV-synchronized acoustic recordings is likely related to the background noise from the equipment, which may degrade signal quality. Perceptual uncertainty in the ground truth (H = 1 vs. H = 2) may further influence model predictions. HSV-synchronized recordings show potential for reliable hoarseness severity estimation even under phonation constraints, with a 500 ms analysis interval sufficient for clinical assessment of voice quality.

Text

Background

Hoarseness caused by different voice disorders may exhibit different acoustic signatures, complicating standardized quantitative assessment. Acoustic voice signals recorded during sustained phonation are commonly used for objective assessment of voice quality [1]. Laryngeal high-speed videoendoscopy (HSV) enables simultaneous recording of acoustic signals during phonation. However, the presence of a rigid endoscope in the oral cavity results in acoustic signals that may differ from natural voice production [2]. Hence, the goals of this study are to determine an optimal, as short as possible, analysis time interval and analyze the potential of HSV-synchronized acoustic signals for machine learning (ML)-based hoarseness severity estimation.

Materials and methods

Two databases containing normal voices, functional and organic voice disorders were constructed. Database D1 comprises 824 HSV-synchronized acoustic recordings of sustained vowel /i/, while Database D2 includes 804 sustained vowel /a/ recordings from speech therapy sessions. Recording segments of 250 ms, 500 ms, and 1000 ms were analyzed. RBH ratings derived from continuous speech served as ground truth [1]. Subjects were categorized into two hoarseness levels (H < 2 vs. H ≥ 2), while ML-based estimated probability was interpreted as a continuous interval-scaled severity score between 0 and 1. Comprehensive acoustic features were extracted and reduced using ensemble feature selection. ML models of varying complexity (logistic regression, SVM, XGBoost, TabNet) were evaluated across durations.

Results

At 1000 ms, Logistic regression for D1 and XGBoost for D2 achieved the best performance, yielding Spearman rank correlations of 0.62 and 0.75 respectively between predicted severity scores and perceptual ratings. Across models, D2 demonstrated consistently higher performance than D1, with mean differences of approximately 10% in accuracy and 8% in ROC-AUC.

Figure 1 [Fig. 1], Figure 2 [Fig. 2]

Figure 1: Model predicted probability scores stratified by ground truth hoarseness ratings H for databases D1 (Logistic regression) and D2 (XGBoost), with Spearman rank correlation (ρ) indicated. A regression line is fitted over predictions.

Figure 2: ROC-AUC score on test sets for D1 and D2 across analysis intervals. Shaded bands represent the standard error.

Conclusion

The lower performance of HSV-synchronized acoustic recordings is likely related to the background noise from the equipment, which may degrade signal quality. Perceptual uncertainty in the ground truth (H = 1 vs. H = 2) may further influence model predictions. Despite these limitations, HSV-synchronized recordings show potential for reliable hoarseness severity estimation even under phonation constraints, with a 500 ms analysis interval sufficient for clinical assessment of voice quality.


References

[1] Dejonckere PH, Bradley P, Clemente P, Cornut G, Crevier-Buchman L, Friedrich G, Van De Heyning P, Remacle M, Woisard V; Committee on Phoniatrics of the European Laryngological Society (ELS). A basic protocol for functional assessment of voice pathology, especially for investigating the efficacy of (phonosurgical) treatments and evaluating new assessment techniques. Guideline elaborated by the Committee on Phoniatrics of the European Laryngological Society (ELS). Eur Arch Otorhinolaryngol. 2001 Feb;258(2):77-82. DOI: 10.1007/s004050000299
[2] Ng ML, Bailey RL. Acoustic changes related to laryngeal examination with a rigid telescope. Folia Phoniatr Logop. 2006;58(5):353-62. DOI: 10.1159/000094569