Journal of Speech, Language, and Hearing Research2026年9月23日

利用音訊與電聲門圖模態自動預測歌唱聲之發聲緊張評分

// Automatic Prediction of Vocal Strain Scores in Singing Voice Using Audio and Electroglottographic Modalities

研究以音訊與電聲門圖(EGG)預測歌唱緊張分數|音訊ρ=0.812、EGGρ=0.759,最佳結合ρ=0.825(皆p<0.001)|EGG提供可解釋的聲門動態並與音訊互補,特徵選擇後之多模態模型具聲樂訓練與臨床監測的客觀評估價值

// The study aimed to predict perceptual strain scores in the singing voice using audio and electroglottographic (EGG) recordings and to evaluate the contributions of each modality, participant metadata, and feature selection. Using a singer-independent train/test split, the authors extracted acoustic (MFCCs, eGeMAPS, wavelet scattering, amplitude modulation), EGG (glottal dynamics), and metadata features, trained regression models (SVR, random forest, ridge) with leave-one-singer-out validation, and applied recursive feature elimination and feature-level fusion. On the held-out test set, audio-only models reached ρ = .812, EGG-only ρ = .759, and the best ridge model combining selected audio features and metadata achieved ρ = .825 (all p < .001). The authors conclude that EGG features are highly interpretable and complementary to audio, and that multimodal modeling with careful feature selection offers a robust objective assessment of singing voice strain supported by demonstrated coupling between glottal and acoustic parameters.

Published 2026年9月22日

RELATED

// 正在尋找相關文獻...