Abstract
Background:
There is a need for fast, accessible, low-cost, and accurate diagnostic methods for early detection of cognitive decline. Dementia diagnoses are usually made years after symptom onset, missing a window of opportunity for early intervention.
Objective:
To evaluate the use of recorded voice features as proxies for cognitive function by using neuropsychological test measures and existing dementia diagnoses.
Methods:
This study analyzed 170 audio recordings, transcripts, and paired neuropsychological test results from 135 participants selected from the Framingham Heart Study (FHS), which includes 97 recordings of cognitively normal participants and 73 recordings of cognitively impaired participants. Acoustic and linguistic features of the voice samples were correlated with cognitive performance measures to verify their association.
Results:
Language and voice features, when combined with demographic variables, performed with an AUC of 0.942 (95% CI 0.929–0.983) in predicting cognitive status. Features with good predictive power included the acoustic features mean spectral slope in the 500–1500 Hz band, variation in the F2 bandwidth, and variation in the Mel-Frequency Cepstral Coefficient (MFCC) 1; the demographic features employment, education, and age; and the text features of number of words, number of compound words, number of unique nouns, and number of proper names.
Conclusion:
Several linguistic and acoustic biomarkers show correlations and predictive power with regard to neuropsychological testing results and cognitive impairment diagnoses, including dementia. This initial study paves the way for a follow-up comprehensive study incorporating the entire FHS cohort.
Keywords
INTRODUCTION
Globally, the number of people living with dementia more than doubled from 1990 to 2016 to 44 million individuals [1]. Currently, almost 6 million Americans live with Alzheimer’s disease (AD), which is expected to rise over 16 million over the next 20–30 years. AD is a type of dementia with symptoms such as memory loss and difficulty with tasks that require coordination, planning, or problem solving [2]. Reductions in functional capacity and mortality from AD are further compounded by lengthy time from onset of mild cognitive impairment (MCI) to AD diagnosis. Final diagnosis can take anywhere from a few weeks to years. Diagnosing AD earlier, at the stage of MCI, would produce $7.9 trillion in cost savings within the United States [1]. The limitations in diagnosing MCI, combined with its mutable nature, have been further implicated as a cause of the heterogeneity in reporting of worldwide MCI incidence [3].
Currently, primary care physicians fail to detect AD in 24% –91% of cases [4–6]. Screening typically relies upon combining cognitive testing with evaluation of functional decline. In cases with complex presentations such as early-onset or mixed-symptom presentations, biomarkers such as tau proteins and amyloid-β42 in cerebrospinal fluid, fluorodeoxyglucose PET, and MRI brain imaging [7] are used. These screening methods suffer from their low accessibility, high cost, invasive nature, and difficulty to detect cognitive decline in the earliest stages [8]. The limitations of our current methods suggest the need to develop fast, accessible, and low-cost diagnostic tools to increase the ease of detecting the presence and progression of cognitive impairment.
Language and voice patterns are affected in the early stages of dementia and may be used as a novel biomarker for cognitive decline [9, 10]. Multiple language pathologies observed in persons with AD include reduced syntactic complexity of language [11, 12], inability to maintain a theme throughout a conversation [13, 14] and perseveration [15, 16]. The changes in language not only co-occur with the disease, but also precede AD, at least one year prior to diagnosis [17], thereby suggesting opportunity for early diagnosis. Voice pathologies observed in persons with AD include alterations in articulation [18], prosody [19–21], and phonologic fluency [22]. Previous work on linking language testing as an indicator of MCI and various degrees of dementia has been primarily focused on small cohorts and used variable testing measures [23].
We aim to validate the use of linguistic, acoustic, and prosodic voice features as proxies for cognitive function using neuropsychological test (NPT) measures. Our goal is to demonstrate methods for automated voice signal processing, feature extraction and selection, and model building to create an interpretable language/voice model of cognitive performance. These methods will serve to build robust linguistic and voice biomarkers for cognitive decline to identify neurodegeneration early, quickly, easily, and at low cost. Here, we conduct feature extraction and statistical analysis of voice samples from an initial sub-sample of 170 NPTs from the Framingham Heart Study (FHS) Cognitive Aging Cohort to identify associations of acoustic and linguistic features with cognitive performance. Our work builds upon a previous analysis of this cohort [24] by roughly doubling its sample size, conducting a robust analysis of FHS neuropsychological test audio quality, and utilizing additional features and feature selection techniques. We report improved classification performance and an error analysis of the trained model.
MATERIALS AND METHODS
Subjects
This study used a subset of NPT with paired audio recordings and clinical characteristics of participants from the FHS [25]. The FHS is a longitudinal, transgenerational cohort study (n≈15,000) which has been running for over 70 years within the United States. More details of FHS data collection, its non-genetic data repository coordinating center, genetic data, and phenotypes can be found in Tables 1 and 2 of [25], in [26], and in [27], respectively. Our focus is on the Cognitive Aging Cohort with audio-recorded NPTs which includes approximately 9,000 individuals followed since 2005.
Cohort characteristics at time of first NPT and each longitudinally recorded clinical cognitive function diagnosis. Some values do not add to 100% because participants’ diagnoses were followed longitudinally and participants had multiple diagnoses over time
Control, unevaluated normal (control group); NC, last date of normal cognition; Impaired, cognitive impairment onset; Mild, mild dementia onset; Moderate, moderate dementia onset; Severe, severe dementia onset. Education: 0 = high school did not graduate, 1 = high school graduate, 2 = some college, 3 = college graduate. Age shows the participants’ age at time of event (NPT or cognitive status diagnosis). NPT 1–4 timelines show the number of years each NPT exam took place after(+) or before(–)first NPT and each cognitive function diagnosis. Demographic values reported as mean (95% confidence interval) unless otherwise stated. Confidence intervals are calculated via bootstrapping. *χ2 test applied to Control, NC, Impaired, Mild, Moderate, and Severe groups. **Characteristic only measured once with no date given for time assessed. ***Age at time of A1 since there is no diagnosis date for pre-evaluated normal cognitive status. †One-way analysis of variance test applied to Control, NC, Impaired, Mild, Moderate, and Severe groups.
Clinical and dementia/AD data
Our study consists of 135 unique participants [age = 73±16, female n = 69(51%) at first NPT] with an average of 1.3 NPTs per participant (170 unique assessments). The sample was hand-selected for another study to ensure that cognitive status was verified for the entire sample and that recordings were usable, and was made available to us for this study. The unit of analysis in the current study is the NPT recording (not the unique participant); because participants are followed longitudinally, a control sample and a case sample could come from the same participant at different points in time. Details of the participant-level clinical characteristics within our study are reported in Table 1, with the timeline of their cognitive status diagnoses and follow up NPTs displayed in Fig. 1.

Participant diagnoses (for those with confirmed dementia progression) and all follow up assessments relative to first neuropsychological assessment/audio recording. The time of first audio recording was mostly during the beginning stages of dementia progression. In fact, the time of first neuropsychological assessment is equal to the Normal Cognition third quartile, Mild Dementia median, and Moderate Dementia first quartile diagnosis dates.
Participants’ cognitive status and the timeline of their dementia progression were determined by an FHS study staff dementia review panel, applying the DSM-IV criteria [2] for dementia and NINCDS–ADRDA [28, 29] for AD. Consensus diagnoses are determined using NPTs and all additional available information, which may include neurology exams, FHS study and external medical records, brain imaging, and for a small subset, interviews with participants’ caregivers and neuropathological confirmation [30]. The control group recordings (n = 97, mean age = 69.5, 45% female) included NPTs for participants who did not undergo dementia review (n = 64, mean age = 62.9, 47% female) as well as NPTs for participants who were determined to be cognitively normal as a result of the dementia review process (n = 33, mean age = 82.2, 42% female). Selection for dementia review is based on whether participants meet the criteria for evidence of cognitive impairment/decline, as previously described [31, 32]. The case group (n = 73, mean age = 83.2, 60% female) included NPTs for participants who had MCI (n = 8, mean age = 76.5, 38% female), mild dementia (n = 40, mean age = 83.7, 58% female), moderate dementia (n = 23, mean age = 84.7, 78% female), and severe dementia (n = 2, mean age = 83.0, 0% female). Detailed information on dementia surveillance within the FHS is reported in the supplementary appendix of Satizabal et al. [30] and elsewhere.
NPT data
The FHS voice samples are taken from recordings of NPT sessions. Each overall assessment is comprised of multiple tests; the ones chosen for this study included Wechsler Memory Scale (WMS) Logical Memory subtests (immediate recall, delayed recall, and recognition), Visual Reproductions testing (immediate recall, delayed recall, and recognition), WMS Paired Associate Learning subtests (immediate recall, delayed recall, and recognition), Wechsler Adult Intelligence Scale (WAIS) Digit Span (forward and backward), the WAIS Similarities subtests, the Verbal Fluency test (naming words beginning with a specified letter (F, A, S), and the Category Naming test (animals)). Full details about the testing protocol can be found in Au et al. [32]. Recordings were made of the entire testing session and we selected recordings from the following tests for analysis: Logical memory (immediate, delayed, and recognition all analyzed as separate tests), Paired Associate Learning (immediate only), Digit Span (forward and backward, which were combined into one test for analysis purposes), Similarities, and Verbal and Category Fluency (fluency tests were also combined into one test for analysis purposes).
Acoustic and linguistic feature extraction
Each NPT was originally recorded as a single audio file that included interviewer and participant audio, with accompanying transcripts. Interviewer speech was removed from each audio file using the timestamps provided in the respective transcript. Then, each recording’s audio file was split into multiple audio files by cognitive assessment task. Audio files without matching NPT data were discarded. The process yielded 1143 individual audio files.
To assess the data utility of the 1143 audio files, a random sample of files (20%) from each available NPT task was assessed by a single reviewer for 1) choppiness and 2) thoroughness of removal of interviewer voice. Choppiness was assessed by calculating the length of the recording divided by the number of removed interviewer audio segments, effectively measuring the length of each uninterrupted span of participant audio. This in part represents the amount of back and forth between the interviewer and the interviewee as determined by the task type; some tasks only ask one question to which a long answer is given, while others have many questions, each of which requires a single word answer. Thus, our metrics for data utility rated tests with lower rates of back and forth dialogue better, which may introduce confounding to our selection of NPTs for analysis below. However, we felt that this potential confounding is not as problematic as poor data utility, since each cut that must be made in the audio file potentially introduces noise, making tasks with much back and forth inherently less suitable for our purposes. The degree of thorough removal of the interviewer’s voice was assessed by the percentage of overall time per recording in which only the participant’s voice was audible (with any remaining percentage reflecting inadequately removed, lingering interviewer voice). Audio recording quality was assessed by calculating the signal to noise ratio over the entire tasks’ audio files for the Logical Memory (Delayed Recall) NPT task only using the snr function in MATLAB (The Math Works Inc., Natick Massachusetts). No other task’s audio quality was assessed since only Logical Memory (Delayed Recall) was chosen for analysis as described in Results.
There are several approaches to extracting acoustic features from audio recordings. Notable approaches include the extraction of the Librosa [33] feature set (spectral bandwidth, flatness, etc.), the Geneva Minimal Acoustic Parameter Set (GeMAPS) [34] feature set (pitch, jitter, harmonic differences, etc.) and AVEC [35] features (skewness, peak range, rising slope, etc.). In this study, the GeMAPS set was extracted from audio recordings for its parsimony and demonstrated effectiveness at capturing voice characteristics as well as larger features sets, which was the purpose of the original development of the feature set [34]. GeMAPS voice features were extracted using the DigiPsych Voice Analysis Pipeline [36]. It is important to note that classical definitions of jitter and shimmer measurements apply to periodic signals in the time domain (e.g. ones found in sustained vowel phonation) to capture variations in frequency and amplitude of the signal. Our values for jitter and shimmer are calculated without the segmentation of speech, and are a periodic in nature [34]. Thus, Jitter and shimmer do not serve as features of connected speech utterances but rather as overall acoustic descriptors.
Similarly, several approaches for extracting linguistic features, such as patterns in word usage and speech content, are available. Here, The Natural Language Toolkit (NLTK) [37] and SpaCy [38] were used. NLTK is an open-source Python package used for raw text processing, including normalization and tokenization in corpus linguistics, which provides part-of-speech tagging. SpaCy, an open-source Python library for advanced natural language processing, is also used for tokenization.
The word count feature from the SpaCy package was utilized to provide context for other part-of-speech tags. Raw noun count may not indicate how frequently a participant uses nouns, because a participant who speaks more might inherently use more nouns than someone who speaks less. This measure was thus used to calculate the normalized noun, verb, adjective, and pronoun rates, which previous work has discussed as possibly associated with cognitive impairment status [39]. We also calculated the pronoun-proper noun ratio by dividing the number of pronouns by the number of proper nouns. Previous work has demonstrated increased pronoun use in AD [40], a decrease in correct pronoun use, and increase in pronoun to noun ratio in those with AD versus healthy controls [9]. Given pronouns often replace proper nouns [41], we hypothesized that the ratio of pronouns and proper nouns is an indicator of cognitive decline. Finally, as described by Roark et al. [42], we also included the propositional density, also called idea density, and the content density.
Statistical analyses
We employed several statistical techniques in order to identify the relationships between NPT and acoustic features. Significance between different groups (Control, NC, Impaired, Mild, Moderate, and Severe) in our demographics table was computed using the Python Scipy [43] package’s chi squared and one way analysis of variance tests for just three (age, education, sex) categorical and continuous variables, respectively, without adjustment for multiple comparisons. The chi squared test on sex was conducted by comparing the observed proportion of female/male participants in each of the six groups to the observed proportion of female/male participants at first NPT. All modeling was performed using the scikit-learn Python package (version 0.21.3). Classification performance is measured using the area under the (nonparametric) receiver operating characteristic curve (AUC), calculated using scikit-learn’s metrics.roc_auc_score function. The curve is constructed empirically by plotting the true and false positive rates over all 170 predicted scores (between 0 and 1) with respect to their binary ground truth labels, over a range of decision thresholds. 95% confidence intervals for AUCs were calculated by measuring the AUC for models trained and tested on bootstrap samples of the original 170 data points (see the Supplementary Material for the details of the procedure).
Correlation
Correlation coefficients between measures of cognitive performance and the acoustic features of voice samples, as well as the linguistic (speech content) features of the samples, were calculated. The Spearman method was chosen so as to account for non-linear correlations without making assumptions about the distribution of the variables [44].
Predicting cognitive status and NPT results
We developed models to predict binary cognitive impairment status from acoustic, demographic, and text features. The two cognitive impairment status categories predicted were ‘normal’ (controls + normal cognition) and ‘impaired’ (MCI + mild + moderate + severe dementia). We first attempted to reproduce results from Alhanai et al. [24] by applying their experimental setup to our data. For this purpose, we first calculated the number of question marks as well as the segment duration for each participant as described in Alhanai et al. and included them along with GeMAPS acoustic features and the following demographic features: age, sex, employment, and education. The categorical features of employment, education, and sex were one-hot encoded to create numerical features. For example, education was encoded into the four binary features “some high school”, “high school”, “some college”, and “college”, such that the applicable category will have a value of 1 and all others will have a value of 0. In a one-hot encoding where exactly one categorical value applies, one category can be dropped, as a 0 in three of four possible categories unambiguously encodes the fourth; in accordance with Alhanai et al.’s protocol, we dropped the first category for each such encoded variable. Next, we filtered out all features which were not correlated with the binary outcome at a 0.05 significance level. To determine which features are most predictive, we then trained several separate Elastic Net models using the Sci-kit Learn implementation (version 0.20.1) [45]. First, we trained a model on acoustic variables (GeMAPS feature set) only. Second, we trained a model on text features only. Third, we trained a model on the demographics features only. We then trained models on each possible combination of these three sets of features. Additionally, we completed a hyperparameter tuning protocol, automatically selecting the best possible configuration settings for each type of model. Each model was trained using a leave-one-out cross-validation approach in accordance with Alhanai et al.’s approach: for each in a predetermined set of hyper parameter combinations, we trained 170 models (each on 169 samples, predicting the one remaining sample using the trained model), and then selected the hyperparameter combination that produced the best-performing model. We thus obtained one set of hyper parameters α and λ, 170 predictions, and 170 sets of coefficients β that were learned by the model. We report the means of each of the coefficients as well as the proportion of models for which each coefficient was non-zero; if a model calculates a zero coefficient for a feature, it means that the model found that feature to be comparatively unimportant in predicting the outcome.
Next, we made several changes in order to build and improve upon these results. First, we included the “base” values for each of the categorical values, as we hypothesized that the effect of individual categories might not be sufficiently captured if they are not all included; for example, dropping out of high school may have a more severe effect than never graduating from college. Thus, we one-hot encoded employment and education into n numeric (binary) columns where n is the number of possible values in the category (Alhanai et al. used n-1 columns). Second, we included text features extracted from transcripts using NLTK and SpaCy; idea density and content density calculated as per Roark et al.; noun, verb, adjective, and pronoun rate; pronoun/proper noun ratio; and participant height, weight, and BMI. Features with no apparent applicability to our use case, such as letter counts, were excluded. Third, we experimented with different variable selection and shrinkage methods outside of the Elastic Net model [46], including Lasso (least absolute shrinkage and selection operator) regression [33, 34] and LassoLars (a version of Lasso using least angle regression [47]). Lasso models differ from Elastic Net models in that Lasso solely implements the L1 regularization penalty (as opposed to both L1 and L2 in Elastic Net) which is then multiplied by hyperparameter λ (sometimes referred to as α). As hyperparameter λ increases, the regularization penalty is amplified and reduces more variables’ coefficients to zero, which removes them from the model. Thus, a higher λ will cause fewer variables to be selected, aiding interpretability but potentially hurting performance. Each of these models performs feature selection to varying degrees with different caveats. Therefore, evaluating different models affords us flexibility when seeking the highest performance while still balancing interpretability. Fourth, in addition to filtering our features based on correlation p-values, we implemented an iterative feature selection approach to reduce the number of features: we trained models on just the acoustic, just the demographic, and just the text features, then used only the selected features in the subsequent models combining the three categories of features. A detailed account of the procedure can be found in the Supplementary Material.
Finally, we completed an error analysis of our best model, choosing a decision threshold above which to call predictions “positive” (impaired) and below which to call them “negative” (normal cognition) such that precision and recall were maximized. We determined the false positive and false negative predictions at this threshold and investigated the factors that contributed to these incorrect predictions. We also reviewed the transcripts for each of the incorrectly classified participants to get a qualitative impression of their impairment status.
RESULTS
Demographic and cognitive comparisons
Our study sample includes a subset of participants (n = 135) in the FHS whose cognitive status and clinical characteristics were followed longitudinally. Since participants were followed longitudinally, diagnosis dates (n = 256) were available for multiple stages of the 71 participants with confirmed dementia progression: Last known date of normal cognition (NC) (n = 69), impairment onset (n = 61), mild dementia (n = 62), moderate dementia (n = 39), and severe dementia (n = 25). The control group consisted of individuals with normal cognition who had undergone NPT but not dementia review (n = 64, 47%). Significant differences in age (p < < 0.001) and education (p < 0.005) of participants were found across cognitive statuses. Since recordings included multiple recordings from single individuals, a subset of the control group samples that came from single individuals contributed to the finding that the control group individuals were younger than the other groups. They were also the most educated, with all other cognitive statuses having similar education levels. Of the 170 available NPT samples, mean age in the control group was 62.8 (58.7–66.8); the subset of individuals who had undergone dementia review but never progressed past normal cognition (n = 33) had a mean age of 82.2 (79.2–85.2). Of participants with known cognitive status, average age at the time of diagnosis was greater at the time of each progressively more severe diagnosis, reflecting progressive decline without recovery. The time of first NPT is equal to the Normal Cognition third quartile, Mild Dementia median, and Moderate Dementia first quartile diagnosis dates, as seen in Fig. 1. At the time of our first recording, the participants’ most advanced cognitive status was: control, n = 64; evaluated NC, n = 24; impaired, n = 8; mild dementia, n = 28, moderate dementia, n = 11; severe dementia, n = 0. Thus, the large majority (84.5%) of participants with a known cognitive status had mild dementia, MCI, or no impairment at the time of our first audio recording. Notably, FHS selected the sample of participants to initially transcribe and share from their cohort of over 5,000 participants deliberately to include a distribution of different cognitive status. In addition, severe dementia cases included in this sub-sample were selected by FHS staff specifically for their ability to be tested and recorded. Many participants with severe dementia cannot complete testing and therefore have no recordings available.
Acoustic and linguistic feature extraction
We selected audio from the Logical Memory (Delayed Recall) test for all subsequent speech feature extraction and analysis because it had the highest balance of data utility and number of available files overall: files (n = 183), length (mean = 42 seconds±17 s, median = 42 s), length per uninterrupted span (mean = 24±15 s, median = 20 s), and participant audio percentage (mean = 90.1% ±16.3, median = 97.7%). Thirteen audio recordings were excluded because they could not be matched exactly with an NPT date or were duplicates, yielding a final sample of 170 audio recordings, transcripts, and NPT results. Data utility between tests was highly variable and is reported in detail in the Supplementary Material. As seen in Fig. 3A, participants with more severe cognitive impairment expectedly performed worse on Logical Memory (delayed recall) than those with less severe impairment.
Audio data quality of the Logical Memory (Delayed Recall) test was poor as evidenced by a low signal-to-noise ratio (mean = –20.44 dB±4.87 dB, median = –21.35 dB). Noise was greater than the recommended minimum [48] (SNR = +15 dB or greater) and was louder than our signal. Audio files were collected at one of three different sampling rates: 8 kHz (n = 48), 22.05 kHz (n = 8), and 44.1 kHz (n = 124). Up until 2006, all audio files exhibit an 8 kHz sampling rate. Audio files collected from the onset of year 2006 through the end of 2008 exhibit either 8 kHz, 22.05 kHz, or 44.1 kHz sampling rates. From year 2009 and beyond, audio files exhibit a sampling rate of either 22.05 kHz or 44.1 kHz. This suggests a change in speech collection protocol due to improved audio collection devices with 8 kHz sampling rate audio collection fully phased out by 2009.
Correlation
In Fig. 2 1 , Spearman correlation coefficients between acoustic and linguistic variables and NPT variables are shown for the subset of variables selected by Lasso. High variation in loudness, among several other measures, was associated with lower cognitive performance, for example lower scores in the paired associate learning test, as well as longer times required to complete the trails test. Longer times to complete the trails test were also associated with larger mean amplitudes for formant 1, 2, and 3, as well as lower amplitude variation in all three. Higher logical memory (immediate recall) scores tended to be associated with lower variation in spectral flux. For linguistic features, we noted that high levels of pronoun usage (pronoun/proper noun ratio) were generally associated with poor cognitive performance, e.g., lower scores in the Logical Memory tests (both immediate and delayed recall) as well as lower performance for Visual Reproductions (both immediate and delayed recall) and lower scores on the Information test (WAIS). Higher levels of conjunction usage and lower usage of possessive endings were slightly associated with better cognitive performance.

Spearman correlations of Lasso reduced (A) acoustic and (B) linguistic set variables with neuropsychological variables in participant story recall data.
Predicting cognitive status and NPT results
Following the approach outlined by Alhanai et al. exactly, we trained Elastic Net models on our own data set of 170 samples, which is about twice the size of the data previously reported on. The model incorporating demographic, audio, and text variables achieved an area under the receiver-operating characteristic (ROC) curve (AUC) of 0.853 (95% CI 0.828–0.920). (The best model achieved an AUC of 0.862(95% CI 0.828–0.920) and incorporated demographic and text features, but no audio features). In accordance with Alhanai et al.’s protocol, we tested calibration using the Hosmer-Lemeshow test, and found the model to be well calibrated. This model selected the number of question marks and the segment duration, as well as several MFCC variables and variation in jitter. Several demographics features were selected as well, including age, sex, employment, and education. Next, we added our additional features but made no further changes to the training protocol in order to crystallize the effects of the additional features. The model utilizing all three feature sets with this setup yielded an improved AUC. We then employed the adjusted model training approach using an Elastic Net model, which further improved performance for an AUC of ∼0.95; however, with this statistical model, 65 features were selected at least 50% of the time. On the other hand, Lasso Lars was able to achieve comparable classification performance while selecting only 21 features at least 50% of the time;17 features were selected 100% of the time. Even though this model was not the model with the highest classification performance, it still achieved very high performance levels in terms of AUC, while selecting fewer features. To balance classification performance with parsimony, which is desirable for interpretability and to reduce the risk of over fitting, we selected this model for further analysis.
Our best model was thus the Lasso Lars model with the following results, shown in Table 3. The models trained exclusively on audio, demographics, and text features achieved a classification performance of 0.758(95% CI 0.752–0.863), 0.799(95% CI 0.738–0.874), and 0.908(95% CI 0.876–0.951), respectively. Incorporating demographics, audio, and text features resulted in the best performance with an AUC of 0.942 (95% CI 0.928–0.983). Here, the hyper parameter λ was 0.0001, and the maximum number of training iterations was 21. (Note that α is always 1 for Lasso models). A total of 28 features were selected by at least one model, and 17 were selected by all 170 models. The mean coefficients for each of the features are shown in Table 2, ordered by coefficient absolute value (strongest contribution first). Amongst the variables selected by 100% of the models, the audio variables with the strongest contributions were mean spectral slope in the 500–1500 Hz band(trend seen in Fig. 3C) and variation in the F2 bandwidth. Parameters reflecting aspects of mel-frequency cepstral coefficient (mean of MFCC 4 and variation in MFCC 1) are represented as well, as is the unvoiced segment length and the mean F1 frequency. Education and employment features had high contributions, as did height and age. The text features among the variables selected by all 170 models were the number of words per sentence, number of compound words (Fig. 3B), number of unique nouns, number of proper names, and the number of questions.
Variables with non-zero coefficients after feature selection via Lasso Lars, ordered by % Selected and then by mean coefficient β absolute value (strongest contribution first). Models were trained to target binary cognitive status (normal cognition versus impairment)
Classification performance, calibration, and hyper parameters for models trained for each combination of features. Calibration was measured with the Hosmer-Lemeshow statistic (HL) which indicates that a model is well calibrated when HL > 0.05. All models are Lasso Lars models and thus have α= 1

Distribution of three prominent NPT (A), Linguistic (B), and Acoustic (C) features by participant diagnosis at the time of NPT using all 170 available NPT with paired audio from the Logical Memory (Delayed Recall) test. The most advanced diagnoses for each participant at the time of NPT were [(Unevaluated Normal, n = 64); (Normal Cognition, n = 33); (Impaired, n = 8); (Mild Dementia, n = 40), (Moderate Dementia, n = 23); (Severe Dementia, n = 2)].
For our error analysis we chose a threshold of ∼0.52, which maximized the F1-score and accuracy of the classifier. All participants scoring higher than this threshold were classified as impaired, while all others were classified as having normal cognition. At this threshold we observed 94 true negative predictions, 61 true positive predictions, 12 false negative predictions and 3 false positive predictions. Thus, sensitivity for impaired individuals (i.e., recall) was 61/(61 + 12) = 0.84 and specificity to impaired individuals (i.e., precision) was 94/(94 + 3) = 0.95. Manual review of the logical memory (delayed recall) test transcripts revealed that two of the three participants with normal cognition incorrectly classified as having cognitive impairment may either be mislabeled or may have memory difficulties unrelated to dementia; their utterances included “I’ve forgotten it” and “I don’t know anything about it, that story”. The third false positive prediction’s transcript did give the impression of an unimpaired person upon review, yet was classified as impaired by our classifier. The features of this sample that drove the predicted score towards the “impaired” classification were, among others, the participant’s high age (88), high variation in F2 bandwidth, and lack of compound words; the features that drove the predicted score closer to the true label of “normal cognition” were the participant’s comparatively small number of questions, body weight, roughly average number of words per sentence, and education. Out of the twelve false negatives, we chose one example to investigate. Most of these participants had transcripts that could have come from a person with normal cognition, so we focus here on an example that was more clearly from an impaired person by the researchers’ judgment. This participant had an actual status of mild dementia but was mistakenly classified as having normal cognition. Major contributions towards the correct impairment classification were made by the average to low number of unique nouns, high age (85), low mean F1 frequency, and high variation in F2 bandwidth. Considerable contributions towards the incorrect classification of “normal cognition” were made by the very small number of questions, the presence of some higher education, a low mean MFCC 4 (which tends to be somewhat lower for the unimpaired in our data set) and low variation in MFCC 1 (which also tends to be somewhat lower for the unimpaired in our data set).
DISCUSSION
The recent scoping review [23] of connected speech assessment in early detection of MCI and AD provides a helpful summary of work to date. Our work here is intended to utilize data much like those shared in the review to automate and provide a linguistic and voice biomarker for cognition. The scoping review found speech production, fluency and semantic outcome measures were best in differentiating participants with MCI from controls. Fluency and semantic measures were best in differentiating mild AD from controls. In addition to fluency and semantic measures, syntactic measures were useful for differentiating moderate AD from controls. Finally, speech production, fluency, and semantic measures were important in differentiating severe AD from controls. This study highlights several shortcomings we hope to overcome using the FHS Cognitive Aging cohort, doubling the size and expanding the analyses conducted within Alhanai et al. [24]. Specifically, previous work was found to be non-specific in terms of labeling MCI as amnestic or affecting a specific domain, focusing primarily on subjects already having dementia and less on those with normal cognition or MCI, used limited testing measures (relying primarily on picture description tasks), and suffered from small sample sizes (with the overwhelming majority including only 3–40 subjects) [23].
The FHS Cognitive Aging cohort provides the largest known and most comprehensive database of recorded voice with robust cognitive labels including brain imaging and NPT, among many other variables. The primary limitations of these data are the need for manual transcription due to poor audio quality and small sample size, reducing the accuracy of automatic speech transcription. A sub-sample has been transcribed as part of an effort to eventually fully transcribe the over 9,500+ recordings, and was made available to us for this initial exploratory work. Moving forward, the need for manual transcription can be addressed by using minimal collection standards including: 1) higher fidelity microphones validated for automated voice analysis which are now almost ubiquitous even across new consumer-grade mobile devices [49], 2) conducting recordings in a clinical setting with microphone placement at a constant, reported distance close to the speaker’s mouth, and 3) protocols that generate longer duration, uninterrupted speech by participants to reduce the need for speaker diarization. Automated speech recognition (ASR) is a possible alternative to manual transcription, as it has been demonstrated to match the accuracy achieved by human transcribers under the right circumstances [50–52]. Our results demonstrate the feasibility of automating voice analysis—on audio collected in controlled environments with data quality in mind—to build a cognitive biomarker. As ASR and audio recording systems continue to improve, automated voice analysis will be possible on audio collected in a wider variety of environments.
Choice of voice samples
For this study, we separated audio recordings into participant voice samples based on cognitive tests, e.g., we created separate voice samples for the logical memory tests and the paired associate learning tests. We found that the quality varied considerably between these voice samples and that the samples of participant voice taken from the Logical Memory test (specifically, the delayed recall version) contained the largest portion of high-quality audio with the Paired Associate Learning (immediate) test having the worst data quality (Supplementary Table 4).
Roark et al.’s previous investigation of language derived measures such as words per clause and content density [42] showed that measures of free language (such as responses from the delayed logical memory task) were more clearly differentiated between individuals with and without MCI than measures of less free language. Our decision to use the logical memory test audio (based on the voice sample quality) for the different cognitive assessment tasks may therefore be justified on the basis of voice sample representativeness as well. We are therefore confident in our choice to focus on voice samples from the delayed story recall task, despite the decreased total length of audio per participant.
Correlations
The correlation coefficients shown in Fig. 2 demonstrate several strong correlations. The voice measures are taken from each participant’s Logical Memory (delayed recall) subtest and correlations to each NPT measure are shown. We found that the mean frequency of the first formant in particular correlated with test scores in a way that shows association with normal cognition; higher F1 frequency is, for example, positively correlated (0.24) with the Boston Naming Test score (more named items correspond to better cognition) and negatively correlated (–0.22) with Time to complete trails A (lower cognitive performance corresponds to a longer period of time required to complete the task). On the other hand, the variation (standard deviation) in formant frequency was inversely associated with cognitive performance, i.e., higher cognitive performance test scores were on average associated with less variation in frequency. This may be explained by the nature of formants as a reflection of the resonance of the vocal tract. If a speaker is intelligible with clear speech, each vowel they speak will produce a series of resonances for which each amplitude peak is captured as a formant. If cognition is declining, speech may be less controlled and formant frequencies may be smaller or less clearly separated: previous work reported that individuals with AD had less control of airflow and thus a more tremulous voice than cognitively unimpaired study participants [53]. Trails A reflects cognitive processing speed and executive function, which, if intact, may correlate with intelligible speech. Time to complete trails B also had strong correlations with several of the same acoustic features indicating some reliability of this association. The finger tapping test was another test with several strong correlations with acoustic features. Similarly, finger tapping is a measure primarily of motor speed, which may tie closely to performance of the vocal cords and changes in formant frequency variation as seen in dysarthric speech [54]. Unvoiced segment length mean and variation both showed a strong negative correlation with Time to complete trails A, which is the opposite association found in other work that shows an increase in speech pauses in those with more advanced dementia [55]. Unvoiced segments have shown ability to aid in discriminating dementia using automated speech analysis [10].
Statistical modeling
Our prediction experiments revealed the utility of voice features, consisting of acoustic and linguistic features, to predict binary cognitive status (normal cognition versus any level of impairment). For parsimony and feature selection, we employed Elastic Net, a comparatively simple linear machine learning model which also allows us to identify the features with the highest impact on classification performance.
Our first important finding is that the approach employed by previous work did not yield comparably high classification performance on our dataset. This is somewhat unexpected, considering that our data set is larger and most likely a superset of this previously analyzed dataset. Additionally, considering the data quality issues uncovered by our audio quality and utility analyses, we believe that our analysis was less limited by noise in the data than previous analysis on this cohort [24]. Possible reasons for the discrepancy in classification performance include that the previously used data subset 1) allowed for training on ample non-random noise/artifact that has no physiological connection to cognition yet significant predictive power to cognition within the dataset (e.g., as cognition worsens, diarization worsens due to increased interviewer interruption), 2) was less diverse than our data and thus easier to model, and 3) that the previously trained models overfit the data.
After applying additional preprocessing steps and an improved training approach, our best Lasso Lars model, which we consider our best model, achieved an AUC of 0.942 (95% CI 0.929–0.983) with 21 features selected at least 50% of the time (17 features were selected 100% of the time). This is a statistically significant improvement over the previously reported model, which reportedly achieved 0.92 on the AUC scale (and an improvement of ∼0.1 over our own execution of Alhanai et al.’s protocol, which demonstrated an AUC of 0.853 (95% CI 0.828–0.920)). The features that were not assigned zero coefficients after training the model were in part reported to be relevant by previous work; for example, Meilán et al. [53] determined that lower levels of pitch modulation were associated with the presence of dementia; Horley et al. [19] found that differences in pitch modulation (variation/std. dev. of F0) can be seen between individuals with and without AD. Our findings provide further evidence for this finding, as our model selected the extent of the interquintile range (from the 20th percentile to the 80th percentile) of the fundamental frequency (i.e. pitch) as a feature negatively associated with cognitive impairment. In other words, the classifier considers a comparatively small variation in pitch to be evidence of cognitive impairment. Additionally, several parameters describing Mel-frequency cepstral coefficients (MFCCs) were also selected by our models, including the mean of and variation in overall MFCC 1 as well as the mean MFCC 4 for only voiced regions. This agrees with a report by Fraser et al. [56], which details that in factor analysis for predicting AD diagnosis, one significant factor was made up of several parameters describing acoustic abnormality, including skewness and kurtosis of a number of MFCCs. Further, MFCCs have been found to be different between groups with and without major depressive disorder, a related mental health pathology; however, MFCC 2 was the only statistically significant contributor [57]. MFCCs, especially the lower order coefficients, may be more important for paralinguistic voice analysis than higher order coefficients, which are more dependent on speech content [34]. Other parameters that have been reported to be significantly different between groups, but which did not provide predictive power across our dataset with our specific choice of model, include shimmer and noise-to-harmonics ratio.
Linguistic features that were retained by our model included words per sentence, number of compound words (as seen in Fig. 3B), number of unique nouns, number of proper names, and number of questions (see Table 2). Of note is that several features that were previously reported in the literature as being associated with cognitive status [9] were not selected by this model. This may be because other features were more representative, or because they were not independent from features already selected, meaning the same information was already represented in other variables. For example, noun and verb rate were not selected; however, the number of words per sentence was selected. It is possible that the information about cognitive performance loss that is captured by noun and verb rate is also captured in the number of words per sentence. Thus, selecting noun rate in addition to words per sentence does not add as much predictive power as selecting some other variable.
Further, several of the selected linguistic features make sense in the context of the specific task the participants are asked to perform: because the logical memory test asks participants to recite details from a story with several unique nouns and proper names (“content units”), the number of remembered story units directly correlates with an increased number of unique nouns and proper names. Similarly, reciting all the details within a few long sentences also reflects that more details are readily remembered. On the other hand, it makes sense that asking more questions of the interviewer, indicating confusion, constitutes evidence of impairment.
The selected demographic features (Table 2) included high school-did not graduate, full-time employment, college education (graduated), height, age, weight, and retirement status. Education, occupation, and age are well documented factors in dementia risk. It is of note that full-time employment strongly depends on age in our data set: 78.8% of participants under 64 were working full-time, while only 3.5% of participants 64 or older worked full-time. While age is also captured in the Age variable, the full-time employment status variable can thus be considered a binary indicator of retirement age or “being elderly”. Presumably on account of the cognitive impairment rate in our data set being 0% under 64, compared to 51% for ages 64 and up, this variable has higher predictive power than the continuous variable for age. Height and weight were selected by our model as well, though BMI was not. Height had a much higher contribution than weight. Though height has no immediately apparent relevance to cognitive status, it appears that it is inversely correlated with age, especially in men.
Additional linguistic analysis
Our findings of increased pronoun use (measured by pronoun to proper noun ratio) at the moderate and severe stages of dementia are similar to previously described findings [9, 58]. Pronoun to proper noun ratios are lower in moderate dementia than in severe dementia. However, the median pronoun to proper noun ratios in normal and MCI individuals are similar and lower than the median pronoun to proper noun ratios of those who have developed dementia. This suggests that cognitive status does not have a significant effect on pronoun to proper noun ratio until the speaker has developed more severe cognitive dysfunction. A reduction in compound words was found to be associated with cognitive decline (Fig. 3B), which has previously been found by other work [59], and contributed to our top features selected by our final model (Table 2).
Limitations
This work has several limitations. First, our choice to target a calculated binary variable describing cognitive status is a limitation. Characteristics of speech may change as individuals develop MCI in a way that is different from participants with advanced dementia. For example, individuals may talk more due to MCI because they begin having difficulty with memory and word finding but retain enough cognitive capacity to “talk around” their limitations. In the FHS interview recordings, we specifically noted that participants may discuss their experience of trying to complete the task, and their difficulties in doing so; they may even reflect on the changes in their cognition over time (“I would have been able to do this a couple of years ago. My memory is just not what it used to be.”). In contrast, individuals experiencing the more advanced stages of dementia may talk significantly less than healthy individuals. In the current study, our comparatively small sample size contributed to our decision to binarize cognitive status; however, this may have contributed to our results being less strong, as the increased speech of MCI may cancel out the decreased speech of moderate and severe dementia in some respects. Our small sample size also limited this study in the evaluation approaches we were able to consider for our predictive models; for example, we were not able to test our model on a held-out test set. Instead, we used leave-one-out cross-validation, which is generally regarded as an acceptable approach for comparatively small data sets. Lastly, we did not adjust in our modeling (e.g., variables of no interest) for the potential confounding posed by our control group being younger and better educated.
Second, we have two separate groups of individuals with normal cognition: controls (not evaluated for dementia) and individuals from the test cohort that were evaluated for dementia but were classified as normal by the Framingham Heart Study staff. Those participants that came to be reviewed were suspected of having cognitive impairment through FHS surveillance methods. It is sensible that only individuals with reasonable suspicion of impairment are reviewed; however, this means that our pre-evaluation group is different from our cognitively intact group, with the intact group perhaps being imperfectly representative of truly intact individuals. Together, they are likely representative of the previously documented broad range of cognitive performance that is considered normal.
A third limitation is the audio utility and quality of the recordings, and their variability over time. The earliest audio recordings suffer from less deliberate data collection and a lower sample rate. Over time, FHS began to record audio with subsequent analysis in mind and made efforts to improve collection standardization and data quality. Given that speech is sampled between 300 Hz and 3.4 kHz in telephony applications, all audio sample rates used in this study are sufficient for capturing speech. However, the 8 kHz sampling rate observed in a minority of audio files allows at most 4 kHz bandwidth, leaving very little margin between the speech frequency band and the sampling bandwidth, which can result in low-fidelity audio. The ideal sample rate to preserve natural speech is 16 kHz as suggested in [60]. However, the quality of audio (e.g., signal-to-noise ratio) is still overall below the recommended minimum standards for acoustic analysis, requiring manual transcription and meticulous verification of the results produced via automatic feature extraction. Thus, our acoustic results must be interpreted with caution in light of acoustic data quality. This limitation is mitigated to some extent by the minimal role acoustic features play in our predictive modeling. As seen in Table 3, the acoustic only model performed poorly and the addition of acoustic data only increased the AUC by 0.0067 when added to our Dem + Highly selected text model. Inclusion of a standardized measure of vocal quality similarity, such as auto-correlation on a ‘sustained phonation’ task, is important for future data collection and study design to ensure that the longitudinal samples per individual are of similar quality. Due to differing collection environments, sample rates, and other changing factors in data collection over time, it was difficult for us to draw a baseline comparison of the audio samples from those individuals with longitudinal voice recording.
Another limitation is the potential specificity of our selected features to the particular neuropsychological testing task we utilized. For example, the number of proper nouns used by a participant is predictive of their cognitive status, but this may in part be because the task was designed to include “content units” (several of which are proper names) such that recalling more content units indicates less impairment. The applicability of this feature may break down when employed in a different context. This represents a general limitation of using tests that are not independent of the outcome label in order to predict outcome; here, the diagnosis is formed in part from the results of neuropsychological testing. While our results should thus be interpreted with caution, we believe that our chosen features include enough information that is independent of the diagnosis criteria to demonstrate the potential of these markers and to serve as a basis of future work which might eliminate the need for NPT altogether, e.g., by utilizing recordings of spontaneous speech.
Finally, in our reproduction of previous results, we were unable to calculate all the input features used by Alhanai et al. from our own data; where this was the case, we used the features that were selected by the previously reported Elastic Net model. However, the selected features may differ for different data sets and depending on the other features available within the data set, so our results unfortunately cannot be considered a true reproduction of results.
Future work
In addition to the digital biomarkers discussed here, the Framingham Heart Study also collects brain imaging data and other clinical measures, including blood amyloid markers. We would like to incorporate these into our analysis in order to validate language/voice biomarkers beyond correlations to NPT measures. We hypothesize voice signals may correlate well with brain imaging when looking at specific regional volumes where atrophy correlates with clinical symptoms (e.g., left anterior medial temporal lobe atrophy with loss of semantic fluency). Further work to validate voice features with large data sets could allow voice to be used as a marker for brain atrophy and specific patterns of neurodegeneration correlating with clinical counterparts (e.g., AD versus FTD). We aim to use the entirety of the FHS Cognitive Aging data set in future work to perform a comprehensive analysis of voice features across the over 5,000 participants for whom data has been collected. We believe this cohort provides the opportunity to build robust language and voice biomarkers for cognition. Additional data would also help address some of the limitations of this work, such as categorizing participants into broad groups due to sample size limitations. Manifestations of cognitive impairment and dementia can differ significantly across stages, and do not always progress in a linear fashion; thus, analyzing individuals in each stage separately has the potential to yield significant insights.
With more data at equal or higher dimensionality, feature selection will remain an important and difficult challenge [61]. In order to avoid overfitting of data [62], a deeper analysis of features will be required using assorted feature selection techniques [63]. Here, we have employed some basic feature selection approaches, including discarding features with low correlation with the outcome and utilizing machine learning models that perform automatic feature selection, such as Elastic Net. However, as the data set grows, more sophisticated and robust feature engineering techniques will have to be considered for future work.
Conclusion
In this study, we have identified associations between NPT scores and characteristics of speech, including acoustic, linguistic, and prosodic features, in a subset of the FHS cohort and demonstrated the utility of these characteristics in cognitive impairment status classification. We also assessed FHS NPT audio quality and demonstrated the potential of these features of speech to predict cognitive status. Thus, this work constitutes important foundational work for future use of FHS data in further studies of voice sample analysis for cognitive impairment classification. Given that the potential for development of linguistic and voice biomarkers is considerable and the impact on disease very strong, funding and conducting further research in this space is crucial.
Footnotes
ACKNOWLEDGMENTS
This work was partially supported by the National Library of Medicine Training Grant (T15LM007442; Authors JAT and HAB), Framingham Heart Study’s National Heart, Lung, and Blood Institute contract (N01-HC-25195; HHSN268201500001I), and NIH grants from the National Institute on Aging (AG008122, AG016495, AG033040) and Defense Advanced Research Projects Agency FA8750-16-C-0299.
