Benchmarking commercial speech recognition and multimodal large language models on dysarthric speech : severity-stratified baselines and architecture-specific prompting effects

Alsayegh, Ali and Masood, Tariq (2026) Benchmarking commercial speech recognition and multimodal large language models on dysarthric speech : severity-stratified baselines and architecture-specific prompting effects. International Journal of Intelligent Systems. ISSN 1098-111X (https://doi.org/10.1155/int/6065038)

[thumbnail of Alsayegh-Masood-IJIS-2026-Benchmarking-commercial-speech-recognition-and-multimodal-large-language-models]
Preview
Text. Filename: Alsayegh-Masood-IJIS-2026-Benchmarking-commercial-speech-recognition-and-multimodal-large-language-models.pdf
Final Published Version
License: Creative Commons Attribution 4.0 logo

Download (1MB)| Preview

Abstract

Voice-based human-machine interaction has become a primary means of accessing intelligent systems, yet individuals with dysarthria are systematically excluded by persistent gaps in recognition accuracy. Whilst automatic speech recognition (ASR) achieves word error rates (WER) below 5% on typical speech, performance degrades sharply for dysarthric speakers; and although multimodal large language models (MLLMs) might compensate for acoustic degradation through contextual reasoning, their zero-shot behaviour on such speech remains uncharacterised. To address this, we evaluate eight commercial speech-to-text services on the TORGO dysarthric speech corpus, comprising four conventional ASR systems (AssemblyAI, Whisper large-v3, Deepgram Nova-3, Nova-3 Medical) and four MLLM-based systems (GPT-4o, GPT-4o Mini, Gemini 2.5 Pro, Gemini 2.5 Flash), across lexical accuracy, semantic preservation, and cost-latency trade-offs. Recognition degraded consistently with severity: mild dysarthria reached low single-digit WER, around 1 to 2% for the leading systems, approaching typical-speech benchmarks, whereas severe dysarthria exceeded 51% WER for every system, with the MLLMs offering no advantage over conventional ASR at this tier under default settings. A four-condition prompt ablation revealed strongly architecture-specific effects, whereby, for the OpenAI multimodal models, any verbatim-transcription prompt reduced severe-tier WER mainly by suppressing non-target-language drift, lowering GPT-4o from 60.1% to 52.9% and GPT-4o Mini from 66.0% to roughly 55%, with the three prompt variants performing similarly; the Gemini models, which did not exhibit this failure mode to the same extent, showed no consistent benefit and sometimes degraded. The semantic metrics correlated strongly with WER and were largely redundant with it in aggregate, but isolated a high-error tail in which communicative intent was partly preserved despite poor lexical accuracy. We offer these severity-stratified, per-speaker baselines as a reusable reference for evidence-based technology selection in assistive voice interface deployment, rather than as a claim of system superiority.

ORCID iDs

Alsayegh, Ali ORCID logoORCID: https://orcid.org/0000-0001-7083-3639 and Masood, Tariq ORCID logoORCID: https://orcid.org/0000-0002-9933-6940;