Neurodegenerative diseases such as Alzheimer’s disease (AD) and frontotemporal dementia (FTD) affect cognition, behavior, and daily functioning. Early diagnosis is critical to reduce their personal, clinical, and social consequences, yet diagnosis and follow-up often rely on costly, time-consuming assessments. Automated speech and language analysis (ASLA) offers a practical window into these disorders, providing fast, objective, and low-cost markers.
Previous work from our team has shown that ASLA can distinguish AD, FTD, and related syndromes; support cognitive phenotyping; and reveal language-based alterations linked to specific cognitive domains and neurophysiological markers. Together, these studies suggest that digital language markers capture clinically meaningful signals.
However, further validation across sociodemographic heterogeneity is needed. Differences in age, gender, education, dialect, multilingualism, and social determinants of health modulate both language production and dementia presentation. Standard group matching reduces average differences between groups but may miss individual variability, while machine-learning splits can reintroduce imbalance.
Our work in 1689 Latin American participants with AD and FTD models this influence using residualization, while differentiating groups and predicting cognitive domains. This approach tests whether ASLA biomarkers remain robust, fair, and clinically informative across heterogeneous populations.