Benchmarking commercial speech recognition and multimodal large language models on dysarthric speech : severity-stratified baselines and architecture-specific prompting effects
Alsayegh, Ali and Masood, Tariq (2026) Benchmarking commercial speech recognition and multimodal large language models on dysarthric speech : severity-stratified baselines and architecture-specific prompting effects. International Journal of Intelligent Systems. ISSN 1098-111X (https://doi.org/10.1155/int/6065038)
Preview |
Text.
Filename: Alsayegh-Masood-IJIS-2026-Benchmarking-commercial-speech-recognition-and-multimodal-large-language-models.pdf
Final Published Version License:
Download (1MB)| Preview |
Abstract
Voice-based human-machine interaction has become a primary means of accessing intelligent systems, yet individuals with dysarthria are systematically excluded by persistent gaps in recognition accuracy. Whilst automatic speech recognition (ASR) achieves word error rates (WER) below 5% on typical speech, performance degrades sharply for dysarthric speakers; and although multimodal large language models (MLLMs) might compensate for acoustic degradation through contextual reasoning, their zero-shot behaviour on such speech remains uncharacterised. To address this, we evaluate eight commercial speech-to-text services on the TORGO dysarthric speech corpus, comprising four conventional ASR systems (AssemblyAI, Whisper large-v3, Deepgram Nova-3, Nova-3 Medical) and four MLLM-based systems (GPT-4o, GPT-4o Mini, Gemini 2.5 Pro, Gemini 2.5 Flash), across lexical accuracy, semantic preservation, and cost-latency trade-offs. Recognition degraded consistently with severity: mild dysarthria reached low single-digit WER, around 1 to 2% for the leading systems, approaching typical-speech benchmarks, whereas severe dysarthria exceeded 51% WER for every system, with the MLLMs offering no advantage over conventional ASR at this tier under default settings. A four-condition prompt ablation revealed strongly architecture-specific effects, whereby, for the OpenAI multimodal models, any verbatim-transcription prompt reduced severe-tier WER mainly by suppressing non-target-language drift, lowering GPT-4o from 60.1% to 52.9% and GPT-4o Mini from 66.0% to roughly 55%, with the three prompt variants performing similarly; the Gemini models, which did not exhibit this failure mode to the same extent, showed no consistent benefit and sometimes degraded. The semantic metrics correlated strongly with WER and were largely redundant with it in aggregate, but isolated a high-error tail in which communicative intent was partly preserved despite poor lexical accuracy. We offer these severity-stratified, per-speaker baselines as a reusable reference for evidence-based technology selection in assistive voice interface deployment, rather than as a claim of system superiority.
ORCID iDs
Alsayegh, Ali
ORCID: https://orcid.org/0000-0001-7083-3639 and Masood, Tariq
ORCID: https://orcid.org/0000-0002-9933-6940;
-
-
Item type: Article ID code: 96698 Dates: DateEvent5 August 2026Published6 July 2026Accepted27 January 2026SubmittedSubjects: Medicine > Internal medicine > Neuroscience. Biological psychiatry. Neuropsychiatry > Communicative disorders. Speech and language disorders
Science > Mathematics > Electronic computers. Computer scienceDepartment: Faculty of Engineering > Design, Manufacture and Engineering Management Depositing user: Pure Administrator Date deposited: 06 Jul 2026 15:19 Last modified: 05 Aug 2026 16:05 Related URLs: URI: https://strathprints.strath.ac.uk/id/eprint/96698
Tools
Tools






