SpeechDx is a 3-year multinational, multilingual observational study (n=2,006) across sites in the US, Australia, and Spain, designed to build a longitudinal speech dataset — linked to blood biomarkers, MRI, neuropsychological testing, and genetic data — to develop and validate speech-based prognostic tools for Alzheimer's disease (AD) and related dementias across the full cognitive spectrum (cognitively normal through AD dementia), with enrichment for preclinical populations.
This is a study design and perspective paper; no clinical outcome results are reported yet. The dataset is projected to generate ~4,000 hours of multilingual (English, Spanish, Catalan) speech recordings with quarterly sampling over 3 years, linked to harmonized clinical and biomarker data across sites.
- No predictive model performance or clinical outcome data are presented; the study is a design framework paper only. - Multilingual design and minimal inclusion criteria introduce linguistic and phenotypic heterogeneity that must be addressed in downstream model development. - Remote longitudinal data collection may be affected by digital literacy gaps, environmental noise, and device-related barriers, potentially introducing algorithmic bias.
Speech-based digital markers are not yet validated for clinical use in AD prognosis, but SpeechDx is building the large, longitudinal, biomarker-linked dataset needed to get there. Clinicians should watch for future validated speech tools as complements to blood biomarkers — especially for predicting who among amyloid-positive asymptomatic individuals will progress to symptomatic disease.
Explore related topics