This diagnostic study enrolled 966 English-speaking adults aged 55+ without prior dementia or MCI from primary care practices in New York City (n=787) and an external validation cohort in Chicago (n=179), recorded during routine visits from August 2020 to December 2021. Machine learning (ML) classifiers were trained on acoustic features extracted from 30-second speech segments using foundation models (Whisper, HuBERT, wav2vec 2.0) and expert-defined methods (eGeMAPS, prosody) to identify cognitive impairment (CI), defined as MoCA score ≥1 SD below age- and education-adjusted norms.
Whisper-derived acoustic features performed best: AUROC 0.733 (95% CI 0.714–0.752) in the primary cohort, with similar results in the external validation cohort (AUROC 0.727; 95% CI 0.714–0.740). As a screening tool on the holdout set, the algorithm achieved sensitivity 68.2% (95% CI 61.8%–74.6%), specificity 63.6% (95% CI 59.8%–67.4%), and positive predictive value 30.4% (95% CI 28.7%–32.1%) against a 21% CI prevalence. Pitch, timing, and speech variability were the top acoustic predictors.
The low PPV (30.4%) reflects a high false-positive rate unsuitable for standalone diagnosis. The study enrolled only English-speaking patients, limiting generalizability to other languages. CI was defined by a single MoCA threshold rather than comprehensive neuropsychological evaluation, which may misclassify borderline cases.
Passive acoustic analysis of routine primary care conversations shows feasibility as a low-burden CI screening aid, but its PPV of ~30% means most positive screens will be false positives — it should trigger, not replace, formal cognitive assessment. Clinicians should not act on this tool alone pending further validation.
Explore related topics