Coming Soon

Analyzing Coronavirus Cough Sounds and Alzheimer's Speech with Deep Learning Models to Discover Common Patterns

A shared 2D audio-analysis framework revealed recurring representation-level proximity between Alzheimer's speech and COVID-19 cough samples, while a weighted CNN ensemble achieved 73% accuracy and 72% macro-F1 for Alzheimer's disease detection.

Conference

IEEE EMBC 2026

Toronto, Canada

Best ensemble

73% / 72%

Accuracy and macro-F1

Status

Paper coming soon

Full text will be linked later

Introduction

Alzheimer's disease (AD) is a progressive neurodegenerative disorder that affects memory, language, communication, and decision-making. Changes in speech can therefore provide useful non-invasive indicators of cognitive impairment.

COVID-19 primarily affects the respiratory system, but neurological and cognitive symptoms have also been reported. Previous research has identified possible biological and clinical relationships between COVID-19 and neurodegenerative disease.

Research Gap

Most previous studies examining relationships between AD and COVID-19 have relied on clinical, genetic, or molecular evidence. Comparatively little research has investigated whether audio recordings from the two conditions exhibit similarities within learned feature representations.

Audio-based analysis may provide a scalable and comparatively accessible way to investigate disease-associated patterns. However, speech and cough recordings differ substantially in content, demographics, and recording conditions, making any cross-dataset findings exploratory rather than diagnostic.

Objective

Primary objective: investigate whether 2D audio representations reveal recurring feature-space patterns between Alzheimer's speech and COVID-19 cough recordings, while also evaluating their effectiveness for speech-based Alzheimer's disease classification.

Specific objectives:

  1. Convert AD speech and COVID-19 cough recordings into seven common 2D audio representations.
  2. Examine cross-cohort representation patterns using CNN feature extraction, dimensionality reduction, and clustering.
  3. Compare VGG-16, VGG-19, and Inception-V3 for AD-versus-control classification.
  4. Determine whether weighted ensemble learning improves performance over the strongest individual CNN model.

Methods

Datasets and Preprocessing

The analysis used 166 ADReSSo recordings: 87 from participants with AD and 79 from cognitively normal controls. Interview recordings were diarized to retain patient-only speech, producing 1,199 AD speech segments and 1,068 control speech segments.

The combined Coswara and Virufy datasets included 1,502 participants: 114 PCR-positive and 1,388 negative. Recordings were divided into non-overlapping individual cough segments to reduce class imbalance.

Audio Representations

Each speech or cough segment was converted into seven 2D representations:

ChromaMel-spectrogramMFCCPower spectrumRaw waveformSpectrogramTonal

Task 1: Cross-Cohort Clustering

VGG-16 was used to extract features from the 2D representations. Principal component analysis reduced feature dimensionality, the elbow method selected the number of clusters, and K-means grouped AD and COVID-19 samples according to their learned representations.

Task 2: Alzheimer's Classification

VGG-16, VGG-19, and Inception-V3 were evaluated across all seven audio representations for binary AD-versus-control classification. Training incorporated validation-based early stopping, class-weighted losses, focal losses, and targeted minority-class augmentation. The strongest representation-specific models were subsequently combined using weighted-average ensemble learning.

Analysis workflow for the cross-cohort clustering task
Figure 1. Task 1 — Cross-Cohort Clustering
Analysis workflow for the Alzheimer's disease classification task
Figure 2. Task 2 — Alzheimer's Classification

Results

Cross-Cohort Clustering

The elbow method selected three to four clusters across the audio encodings. COVID-19 participants in their early 30s and AD participants around age 70 repeatedly co-occurred in learned-feature clusters.

This represents exploratory feature-space proximity, not evidence of causation, disease equivalence, or shared cognitive impairment.

Table 1: Cross-Cohort Cluster Composition

For each audio representation, the K-means cluster with AD and COVID-19 proportions closest to 50/50 is shown; ages are reported as mean ± SD.

RepresentationClusternAD%COVID%Mean Age: AD / COVID-19
Chroma1/3124653570 ± 6.5 / 33 ± 12.2
PowerSPEC3/3121118970 ± 6.2 / 33 ± 13.2
Raw2/3124703070 ± 6.8 / 32 ± 15.9
MFCC2/3124138769 ± 6.0 / 33 ± 13.2
SPEC3/4145435769 ± 6.8 / 34 ± 14.0
Tonal1/3173465470 ± 6.6 / 33 ± 12.5
Mel-spectrogram2/4142396171 ± 5.8 / 33 ± 12.5

Alzheimer's Classification Performance

Inception-V3 with MFCC was the strongest individual model, achieving 70% accuracy and 70% macro-F1. Post-hoc five-fold evaluation preserved the overall ranking of the strongest model-representation combinations.

Ensemble Learning

Weighted Average combined the strongest representation-specific models. It improved the best individual model by 3 percentage points in accuracy and 2 percentage points in macro-F1.

Best individual

70% / 70%

Inception-V3 + MFCC

Weighted ensemble

73% / 72%

Accuracy and macro-F1

ApproachAccuracyMacro-F1
Inception-V3 + MFCC70%70%
Weighted Ensemble73%72%

Discussion

The clustering results suggest that the CNN representations capture recurring acoustic structure across the datasets. However, differences in age, recording task, audio modality, and collection conditions make the finding hypothesis-generating rather than clinical evidence.

MFCC was the strongest input for Inception-V3, and weighted averaging improved performance by combining representation-specific predictions. The audio-only CNN approach is simpler than multimodal transformer systems, although some multimodal approaches report higher accuracy.

Limitations

  • Speech and cough are fundamentally different vocal events.
  • Static 2D encodings may discard temporal information.
  • Dataset size, imbalance, and demographic differences limit generalizability.
  • Independent external and clinical validation is required.

Conclusion

  • Shared 2D audio representations supported both exploratory cross-cohort clustering and AD classification.
  • Inception-V3 with MFCC achieved 70% accuracy and 70% macro-F1; weighted ensemble learning increased performance to 73% accuracy and 72% macro-F1.
  • The clustering result remains exploratory and requires validation using larger, demographically matched cohorts and temporal or multimodal models.

References

  1. Luz S, Haider F, de la Fuente S, Fromm D, MacWhinney B. Detecting Cognitive Decline Using Speech Only: The ADReSSo Challenge. Interspeech, 2021.
  2. Mohammed EA, Keyhani M, Sanati-Nezhad A, Hejazi SH, Far BH. An Ensemble Learning Approach to Digital Corona Virus Preliminary Screening from Cough Sounds. Scientific Reports, 2021.
  3. Szegedy C, Vanhoucke V, Ioffe S, Shlens J, Wojna Z. Rethinking the Inception Architecture for Computer Vision. CVPR, 2016.

Acknowledgements

Supported by Hotchkiss Brain Institute, the Natural Sciences and Engineering Research Council of Canada, and the Canada Research Chairs Program. Secondary, de-identified public datasets were used. The authors declare no competing interests.

Emad A. Mohammed, Salma J. Ahmed, Emiliano Garcia Ochoa, Behrouz H. Far, Amir Sanati Nezhad · Physics and Computer Science Department, Wilfrid Laurier University; Thompson Rivers University; University of Calgary · IEEE EMBC 2026 · Toronto, Canada · Paper ID 1313