inklap

Multi-label voice disorder classification using raw waveforms

GÖKAY DİŞKEN · Turkish Journal of Electrical Engineering and Computer Sciences · 2024

Automated voice disorder systems that distinguish pathological voices from healthy ones have been developed with the aid of machine learning methods. Both clinicians and patients can benefit from these systems as they provide many advantages, compared to the invasive techniques. These systems can produce binary (healthy/pathological) or multi-class (healthy/selected pathologies) decisions. However, multiple disorders might exist in an individual’s voice. Multi-label classification should be considered in such cases. By this time, only a single report is available on this topic, where hand-crafted features were used, and a data augmentation technique was utilized to overcome class imbalances. In this study, a similar experimental setup is followed to investigate the suitability of raw voice signals as inputs for multi-label classification. A deep learning model which consists of residual blocks and a novel gating mechanism is proposed. The gating mechanism weighs the channels of a residual block’s output based on both its output and the previous layer’s output. Using a SincNet filterbank that operates directly on the raw waveform as the initial layer, 0.99 accuracy and 0.98 F1 score

📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً