Abstract
Keywords
Introduction
Cognitive assessment tools are critical to the processes of diagnosis, treatment planning, and on-going evaluations of many populations. Many assessment tools are available to measure cognitive abilities. Given the heterogeneity of individuals with cognitive impairments due to various etiologies (e.g., brain injury, stroke, dementia), clinicians tend to administer broad assessments to assist with identification of areas that require in-depth evaluation using specialized, time-intensive methods. Although brief to administer, these broad assessments or screening tools require significant time to score. Additionally, to reduce therapists’ workload, the tools are sometimes administered by professionals from a wide range of disciplines with different training backgrounds, potentially affecting standardization of assessment administration and introducing error (Hinojosa & Kramer, 2014).
Computerized, self-administered cognitive assessment tools may offer one alternative to current assessment methods. Given the increase in use of tablets and other mobile devices in the provision of clinical services (Fernandes, 2011; Michi, Yamashita, Imai, Suzuki, & Yoshida, 1993; Sackaloo, 2013) as well as by people with disabilities (e.g., Morris, Mueller, Jones, & Lippincott, 2014), a computerized tool that can be implemented on such a platform may be valuable for therapists. Implementation and evaluation of computerized assessment tools has been occurring for the last few decades as computers have become widely available. Now, the development of relatively inexpensive, portable computerized tablets has further increased the utility of these tools (Wild, Howieson, Webbe, Seelye, & Kaye, 2008). For example, 53% of occupational therapists reported the use of tablets to access applications (apps) in therapy (Erickson, 2015). For clinicians, iPads have been found to decrease bias and human error that might affect evaluation (Wild et al., 2008) and have the potential to increase productivity and intervention time.
Advantages to computerized cognitive assessments
Several potential benefits exist for computerized, self-administered cognitive assessment tools as compared to paper and pencil assessments. For example, paper and pencil assessments may require significant set-up and preparation of materials for clinicians. Whereas, computerized, self-administered cognitive assessment tools, eliminate the need for this type of preparation as all stimuli are contained as part of the software. All computerized instructions are provided in a consistent, standardized manner versus the potential variability that may occur when therapists administer paper and pencil assessments (Wild et al., 2008). Additionally, implementation of other health information technology has demonstrated an increase in efficiency (Buntin, Burke, Hoaglin & Blumenthal, 2011). Scores from computerized assessment tools are generated automatically, providing clinicians valuable time to conduct other critical clinical tasks. Finally, computerized assessment tools automatically generate highly accurate information about response time for each task that an individual completes. Speed of response may sometimes be more sensitive to subtle cognitive differences than accuracy (Nicholl et al., 1995; Wild et al., 2008). In addition to the clinical benefits, researchers have identified the utility of computerized cognitive testing to effectively evaluate and monitor cognition in large neuroepidemiological studies that draw from populations across a wide geographical area (Fredrickson et al., 2009; Onoda et al., 2013).
Given the potential advantages of using computerized, cognitive assessments it is not surprising that many such tools are in development and have been evaluated across multiple research studies (Fredrickson et al., 2009; Collie et al., 2003; Onoda et al., 2013; Wild et al., 2008). Recently, studies examining the validity and utility of computerized cognitive assessments have focused on two primary clinical populations: individuals with dementia or who are at risk for dementia, and individuals with concussions.
Computerized cognitive assessment in dementia
Computerized cognitive assessments have been developed and implemented with older individuals at risk for cognitive decline (Fredrickson et al., 2009; Onoda et al., 2013; Wild et al., 2008). For example, Onoda and colleagues (2013) identified strong validity in a computerized cognitive assessment administered via iPad. This tool was designed to screen older adults for the presence of symptoms related to dementia. Similarly, Fredrickson and colleagues (2009) found that a game-like computerized test (i.e., CogState) showed stability and high test-retest reliability in healthy adults over age 50 years who completed testing four times over a one-year period.
To summarize some of the research in this area, two groups of researchers reviewed computerized cognitive tests appropriate for testing elderly adults (Wild et al., 2008; Zygouris & Tsolaki, 2014). Wild and colleagues (2008) reviewed 11 and Zygouris and Tsolaki (2014) reviewed 17 computerized cognitive batteries. Both groups identified that strengths of these batteries are similar to those typically associated with computerized testing: administration and stimuli presentation standardization, automated scoring, accurate response time measures, access to real-time results, and readily available comparisons to prior performances. These strengths should be considered in the context of psychometric properties of the tools and careful examination of the variety of available tools (Wild et al., 2008; Zygouris & Tsolaki, 2014).
Computerized cognitive assessment in concussion
Other cognitive assessment tools have been developed to monitor changes in cognition following concussion. Awareness of the importance of cognitive assessment in this population is increasing as a result of the high incidence of concussions and the consensus statement from the fourth International Conference on Concussion indicating a cognitive evaluation is an essential part of concussion management (McCrory et al., 2013). Computerized assessments, and in particular, self-administered cognitive assessment tools may prove useful for these types of evaluations which may occur outside of regular clinical settings. For example, CogSport was determined to be highly reliable for assessment of cognition in elite athletes and young adults (Collie et al., 2003).
Standardized Touchscreen Assessment of Cognition
Despite potential advantages of computerized cognitive assessments and their frequent use in dementia and concussion management, a need exists to compare each newly developed cognitive assessment tool to those currently used by practicing clinicians. The Standardized Touchscreen Assessment of Cognition (STAC) (Cognitive Innovations LLC, 2016) is one such tool. The STAC differs from other cognitive assessments because it was designed to be appropriate for a wide range of individuals with cognitive deficits and etiologies rather than just for concussion or dementia management. The STAC is a criterion-referenced test that assesses cognitive functions. Similar to frequently used paper and pencil cognitive assessments, the STAC consists of tasks designed to assess attention, memory, and executive function skills. Additionally, the STAC includes subtests that measure linguistic functions (i.e., confrontation naming). The results, which the STAC automatically generates, give clinicians both quantitative and qualitative information about a person’s cognitive abilities. Quantitative information includes accuracy of performance and speed of response. Qualitative information relates to the type and pattern of errors. Different from other paper and pencil tasks, the STAC is self-administered using an iPad application, which is intended to minimize examiner bias and standardize administration (Cognitive Innovations LLC, 2016).
Given the STAC was designed to evaluate cognition in multiple populations of adults with cognitive deficits, it should be compared to cognitive assessments that are similarly appropriate for a wide range of potential cognitive and linguistic impairments resulting from multiple etiologies (e.g., stroke, traumatic brain injury). The Cognitive-Linguistic Quick Test (CLQT) (Helm-Estabrooks, 2001) and the Cognitive Assessment of Minnesota (CAM) (Rustad et al., 1993) are commonly used paper and pencil cognitive assessment tools used in speech-language pathology and occupational therapy respectively. Prior to making comparisons in clinical populations, evaluation of neurotypical adults provides an important foundation for beginning to validate the STAC. Therefore, the purpose of this study was to collect data with neurotypical adults ages 18–85 years old using the STAC and compare these results to those from the CLQT and CAM.
The specific research questions were: What are the correlations between similar tasks (subtests) on the STAC versus the CLQT and CAM? Do people with varied levels of iPad comfort perform differently on the STAC versus the CLQT and CAM? Do some age groups have stronger correlations between STAC versus the CLQT and CAM than other age groups?
Methods
Participants
Eighty-eight adults (42 males, 46 females) without cognitive impairment participated in this study. All participants spoke American English as their primary language and scored at least 25 out of 30 on the Mini-Mental State Exam (MMSE) (Folstein, Folstein, & McHugh, 1975). Six participants were left-handed. The participants were between 18 and 85 years old using a stratified recruitment strategy for each of five age groups: 18–29, 30–45, 46–59, 60–74, and 75–85 years. Four of the age groups included 20 total participants; the 75–85 age group only had 8 participants (two males and six females) due to difficulty with recruitment in this age range. Table 1 provides data regarding the mean age and gender of each age group.
Gender and mean age by participant group
Gender and mean age by participant group
The researchers excluded participants with potential neurological, hearing, or vision impairments based on screening results described below. Each participant provided self-reported information regarding education level, ethnicity, as well as cell phone and tablet usage (Table 2).
Participant demographic information (n = 88)
The researchers conducted participant interviews guided by a demographic form to gain background information regarding each participant. The form consisted of questions pertaining to participants’ age, education level, primary language, history of neurological impairments, handedness, and touch screen use. As part of the demographic form, participants also confirmed that they did not have any hearing, uncorrected vision, or motor impairments that would affect their completion of the experimental tasks. The researchers administered the MMSE (Folstein, Folstein, & McHugh, 1975) to rule out the presence of a major neurological impairment. The MMSE provided further confirmation that hearing, vision, and motor impairments would not interfere with participants’ performance of experimental tasks. Next, the researchers collected data from three other assessment tools: CLQT (Helm-Estabrooks, 2001), CAM (Rustad et al., 1993) and STAC (Cognitive Innovations LLC, 2013).
The CLQT is a criterion-referenced test developed to assess a person’s performance in five cognitive areas including: attention, memory, executive functions, language, and visuospatial skills. It has a total of 10 paper and pencil subtests including both verbal and non-verbal tasks. These subtests include: personal facts, symbol cancellation, confrontation naming, clock drawing, story retelling, symbol trails, generative naming, design memory, mazes, and design generation. It is appropriate for individuals between the ages of 18 and 89 years old and provides criterion cut scores for each cognitive domain relative to two age groups (i.e., 18–69 years; 70–89 years). The CLQT can be administered in 15 to 30 minutes (Helm-Estabrooks, 2001).
The CAM is a criterion-referenced screening test that employs a hierarchical approach to evaluate a range of cognitive skills. The CAM is made up of 17 paper and pencil subtests ranging from simple to complex. It examines a wide variety of skills in the areas of attention, memory, visual neglect, temporal awareness, mathematical computation, and following directions. The CAM is appropriate for adults with acquired brain injury. The associated normative sample included adults ages 18 to 70 years old without neurological impairments. This test may take between 30 to 40 minutes to administer (Rustad et al., 1993).
The STAC is criterion-reference test that is self-administered and provides quantitative and qualitative information about cognitive abilities (Cognitive Innovations LLC, 2016). It includes 15 subtests that assess attention, memory, language, and executive function skills. The STAC tracks response time for most items. The STAC provides both written and verbal directions given in both male and female digitized (i.e., recorded) voices. The STAC was administered on a standard second generation iPad with a screen size of 9.7 inches, and participants were invited to adjust the volume as needed. The average key size on the touch screen keyboard was standard (i.e., about 0.5 by 0.5 inches). As part of the STAC administration, participants identified if they were comfortable, somewhat comfortable, or not at all comfortable using an iPad prior to beginning the assessment. The STAC takes about 15 to 30 minutes to administer for most individuals.
Procedures
Study sessions took place in an outpatient clinic therapy room or in a quiet space in the participants’ home. Potential participants signed an informed consent form approved by the institutional review board. Individuals who provided consent proceeded with the single data collection session. First, the researchers conducted individual interviews using the demographic form as a guide. Then, the researchers administered the MMSE to screen each participant for possible cognitive deficits. After the participant passed the MMSE and the examiner determined he or she met study criteria, the researchers administered each of the assessment tools (i.e., CAM, STAC, and CLQT) in an order that was randomized for each participant. The examiners carefully followed each test’s protocol as instructed by the manual. If necessary, cues were given as instructed by each test manual. Following administration, the researchers scored the assessment tools and entered the data into SPSS version 23.0 for later analysis.
Data analysis
The researchers scored the MMSE, CLQT, and CAM according to the manual. The STAC generated a summary sheet that included accuracy and response time for items completed. The researchers matched subtests according to task similarity and cognitive skills measured resulting in a list of 8 comparisons between the CLQT and STAC, and 10 comparisons between the CAM and STAC (See Table 3 for matched subtests). The researchers used SPSS version 23.0 to conduct correlations among the matched subtests. Correlations were interpreted using the following values: above 0.75 was considered a good to excellent relationship; 0.50 to 0.75 was considered a moderate to good relationship; 0.25 to 0.50 was considered a fair to moderate relationship and less than 0.25 was considered to have little or no relationship (Portney & Watkins, 2000).
Matched subtests
Matched subtests
The researchers identified five areas of assessment tasks to conduct the additional analyses related to iPad comfort and age. Areas selected were those that are frequently assessed in people with cognitive impairments and those which demonstrated variability within the current sample. These areas included: generative naming category, generative naming first letter, immediate auditory memory, delayed auditory memory, and auditory working memory. These analyses were conducted using the iPad and paper and pencil subtests related to each area of assessment. First, the researchers determined the means and standard deviations for each subtest with groups separated by the variable iPad comfort. Then, a one-way multivariate analysis of variance (MANOVA) with post-hoc analyses with Bonferroni correction was conducted to determine significant differences between the three levels of iPad comfort on each task. Finally, the researchers separated the participants into the five assigned age groups and computed descriptive correlations for each age group across the five areas of assessment.
Results for research question one include the means, standard deviations and correlations for the target areas of measurement across the STAC, CAM, and CLQT. Participants performed without any variability on five of the selected subtests (CAM Memory/Orientation: Recent, CAM Object Identification, CAM Line Bisection, CLQT Personal Facts, and CLQT Confrontation Naming). The means and standard deviations for each of selected subtests appear in Table 4.
Means and Standard Deviations for selected tasks on the STAC, CAM, and CLQT
Means and Standard Deviations for selected tasks on the STAC, CAM, and CLQT
*Raw data rather than score used for calculations. **Subscore 1 used for calculations.
The strongest correlations were within the areas of generative naming category and generative naming first letter between the STAC and CLQT (0.70 and 0.64 respectively). Measurements of immediate visual memory were also correlated between the STAC and CLQT (0.53). Weaker correlations were identified between the STAC and CAM with the highest being in areas of visual working memory (0.43), auditory working memory (0.37), delayed visual memory (0.34), and delayed auditory memory (0.31). Correlations could not be computed within the areas of orientation and confrontation naming because performance on both the CAM and CLQT resulted in no variability. Additionally, participants performed without variability within the areas of visual attention and immediate auditory memory on the CAM. Correlations for the selected tasks are displayed in Table 5.
Correlations within selected areas of measurement for the CAM, CLQT, and STAC
The results for research question two include the five areas of measurement selected for follow up evaluation. First, the researchers considered the results across participants’ reported iPad comfort level. Forty-three percent of participants (n = 38) reported they were “comfortable” with an iPad; twenty-seven percent of participants (n = 24) reported that they were “somewhat comfortable” with an iPad; and thirty percent of participants (n = 26) identified that they were “not comfortable” using an iPad. Six subtests (2 areas of measurement) were excluded from this analysis because on at least one of the subtests being compared, the participants performed without variability. For example, within the area of immediate auditory memory all participants performed without variability for the CAM regardless of their iPad comfort level. Additionally, within this same area of assessment (i.e., immediate auditory memory), participants who reported being “somewhat comfortable” with an iPad performed without variability on the STAC.
Findings for the remaining 11 subtests, showed significant differences across the three levels of iPad comfort for 4 subtests. Of these 4 subtests, 3 were from the STAC and one was from the CLQT. The post-hoc analysis indicated that participants reporting being comfortable with iPads scored significantly higher than participants who reported being not at all comfortable with an iPad. See Table 6 for the means, standard deviations, and p values. Table 7 shows the post-hoc analyses for each relevant variable.
Means, standard deviations, and p values for each subtest across levels of iPad comfort
Shading = significant at the p < 0.005 level.
Post-hoc analyses of iPad comfort levels with p values
Shading = significant at the p < 0.005 level.
Finally, correlations among similar subtests were calculated for each of the five age groups. Across the different age groups, the number and strength of correlations varied within different areas of assessment. For the Generative Naming Category Subtests, the STAC and CLQT were at least moderately correlated for all age groups except 30–45 years (r = 0.47). These correlations ranged from 0.47 to 0.70. Similar results were evident from the Generative Naming First Letter Subtests such that the STAC and CLQT were at least moderately correlated for all age groups except 30–45 years (r = 0.09) and 75–85 years (0.42). Aside from the 30–45 years age group, correlations ranged from 0.55 to 0.83 for this area of assessment. Within the Auditory Immediate Memory Subtests, the correlations ranged from 0.12 to 0.80 for the 18–29 years old and the 75–85 years old respectively. Only the 75–85 years age group correlation was at least moderately strong. No correlations within the Auditory Delayed Memory or Auditory Working Memory Subtests reached the level of moderately strong.
This study investigated the validity of the STAC when compared to common paper and pencil assessments of cognition, namely the CLQT and the CAM. Results indicated that the STAC has a fair to good relationship with these tests when looking at a variety of subtests. Furthermore, self-reported comfort level was related to STAC performance with people reporting high levels of comfort with the iPad performing better than people who reported not being at all comfortable with iPad use.
One key finding resulting from this investigation was the relationship observed between comfort level and performance on the STAC. As noted in reviews of computerized cognitive assessments, this element is often overlooked or not discussed in the development or psychometric testing of these types of tools (Wild et al., 2008; Zygouris & Tsolaki, 2014). This finding suggests that mode of delivery has the potential to impact performance on cognitive screenings delivered via touchscreen. This is consistent with findings from Raymond and colleagues (2006) for a computer-based cognitive test, the MicroCog battery, which found improved performance on the second testing time. This improvement was not expected and suggests that the observed change could be due to an increased comfort with technology. While some advocate that if a test has good discriminant validity the response mode (i.e., touchscreen versus paper and pencil) is not as crucial (Zygouris & Tsolaki, 2014), however comfort level could be important in understanding the concurrent validity of some of technology based tools when compared to pen and paper based equivalents. Furthermore, it will be important for clinicians to understand the possibility of decreased performance during a screening if a client is not familiar with technology or touch screen assessments and to interpret these results with caution. That is, people with low levels of iPad comfort may still be able to complete the assessment via iPad; however, their performance on some tasks may be lower than expected from a paper and pencil assessment. This possibility should be evaluated on an on-going basis as the general population becomes increasingly comfortable with touch screen technology in various aspects of daily life.
One major difference between the touchscreen, computerized assessment (i.e., STAC) and the paper and pencil assessment tools (i.e., CLQT and CAM) was that the STAC subtests often required typed responses rather than spoken responses. Given the differential language and motor demands of each response type, it might be expected that this would result in large differences in participants’ performances. However, the Generative Naming tasks which require participants to either say (CLQT) or type (STAC) as many words as they can in a specific time limit, resulted in moderate correlations (0.70 and 0.64 respectively for Category and First Letter Naming) suggesting a relationship between these tasks despite the different modalities used. The STAC allowed the participants more time to complete the task than the CLQT. Additionally, although a relationship was evident, participants on average named more words for the CLQT category generative naming than the matching STAC subtest (23.44 versus 15.18). In contrast, the generative first letter naming mean score was slightly lower for the CLQT subtest compared to the STAC subtest (13.77 versus 14.28). Consideration of these differences is critical to the future use of STAC (which has at least four items with typing) or similar touch screen assessments that rely on typed rather than verbal responses because generative naming is used across multiple populations for evaluation of language and cognitive skills. Specifically, when used as verbal tasks, both first letter and category generative naming have been found to be sensitive to changes related to brain damage. For example, people with the Alzheimer’s Disease tend to demonstrate greater impairments in category versus first letter tasks (Monsch et al., 1992). Similarly, older adults tend to demonstrate poorer performance with category versus first letter naming, although both types of naming may decline with age (e.g., Brickman et al., 2005; Loonstra, Tarlow, & Sellers, 2001; Tombaugh, Kozak, & Rees, 1999). In the future, clinicians may be inclined to rely on assessment tools that use typed rather than spoken measures of language abilities. However, clinicians should carefully consider the norms for each modality because in a neurotypical population these results were correlated but the means were different. That is, clinicians may not expect as many typed responses compared to spoken responses during category generative naming tasks. This warrants further research in clinical populations with careful evaluation of any possible differences across modalities.
The present study allowed the researchers to compare the STAC to frequently administrated cognitive assessment tools. As such, the findings suggest modifications to STAC that would likely improve its measurement of cognitive abilities. The first change related to the use of line drawings for the STAC Confrontation Naming subtest. All participants scored 100% on the CLQT and CAM confrontation naming subtests. However, for the almost identical STAC subtests, 12 of the 88 participants typed wheel for spoke because of confusion about the picture. Changes to this picture would likely result in similar performance for participants without cognitive deficits across the three assessment tools. On other STAC subtests such as the Immediate Recall subtest, the researchers found that spelling errors were sometimes counted as incorrect responses and needed to be rechecked by the examiner. Test developers could consider creating settings for specific subtests that allow minor spelling errors made frequently by individuals without cognitive impairments to be counted as correct. Similarly, given the differences noted above on generative naming tasks, voice recognition software may be considered as another mode of response on the STAC. Finally, on the STAC Attention Executive Function Subtest, which is designed as a trail making task, many participants struggled to complete the task because of the sensitivity of the touch screen and confusion about the directions. Developers of touch screen tasks that require participants to drag their fingers across the screen (e.g., trails type tasks) may want to consider other modifications or changes to the touch screen sensitivity. Additionally, the CLQT Trail Making subtest allows for multiple practice items, which might be helpful on the STAC subtest as well to ensure that clients understand how to complete the task. Taken together, consideration of these types of changes may improve the accuracy of touchscreen based cognitive assessments.
Limitations
The primary study limitation was the limited amount of variability that was present in this sample. While the researchers purposely restricted participation to individuals who were neurotypical for this initial study, this constraint may have led to challenges in comparing the subtests as there were some subtests with little to no variability preventing statistical analysis. This limited variability could also potentially explain the lack of correlations present between age and performance on the STAC. It may have been that the inclusion criteria decreased variability across the sample limiting the impact of age that would have been expected. It is also important to note that the sample lack geographic variability with most of the participants coming from no more than an hour away from a mid-Atlantic urban center. A further limitation related to composition of the sample such that most participants were highly educated (57% had at least a bachelor’s degree) and Caucasian (94% of the sample).
Future research
Future research should focus on establishing the validity of the STAC in populations who experience cognitive deficits to understand the viability of this tool for use in clinical practice. When investigated in clinical populations, it may be advantageous to compare the tool to paper and pencil tools and extend the research to establish the discriminant validity of the STAC as this may avoid a potential impact of iPad comfort on STAC performance. Further research may also include the comparison of performance on the STAC to performance on functional measures of cognition, such as the Executive Function Performance Test (Baum et al., 2008), Assessment of Motor and Process Skills (Fisher & Jones, 1999), and the Test of Everyday Attention (Robertson, Ward, Ridgeway, Nimmo-Smith, & McAnespie, 1991) to see if it is predictive of performance on function based tests.
Conclusion
Overall, the findings from this preliminary study suggest that while some areas of assessment using the STAC likely provide valid results, other areas did not correlate well with currently used paper and pencil cognitive assessment. Additionally, comfort with the iPad and modality of response (i.e., spoken versus typed responses) appeared to affect performance in this neurotypical population. However, these results should be confirmed in a clinical population with cognitive impairments who would likely present with greater variability across areas of assessment.
Conflict of interest
Cognitive Innovations, Inc. provided participant incentives and copies of the STAC application to support this project. The authors have no other relevant financial or non-financial disclosures.
Footnotes
Acknowledgments
We thank our participants for their time and contributions to this project. We also want to thank Heather Coles and Simon Carson for their support in completing this project.
