Abstract
Students with complex communication needs have increasingly been using non-dedicated communication systems, such as mobile devices, to support their communication needs. This in turn, has led to an increased used of augmentative and alternative communication apps. The main challenge currently faced is the lack of empirically validated apps and evaluation systems to assess the features of the apps. As a result, this study attempted to determine the reliability of an app evaluation tool that was grounded in the components of the feature matching model. The goal was also to identify if the app evaluation tool could be used to evaluate various types of augmentative and alternative communication apps. Participants evaluated apps across the dimensions of usability, output, and display. Results suggest that expert raters were more reliability than novice raters across the various types of apps. Practical implications and future research directions are discussed.
Keywords
Examining the Reliability of an AAC App Evaluation Tool: Differences Between Expert and Novice Raters
Over five million people are estimated to present challenges in effectively communicating their wants and needs (Beukelman & Light, 2020). Individuals with complex communication needs benefit from the use of augmentative and alternative communication (AAC). These supports can help develop the communicative competence of those with such needs, which in turn, can facilitate greater independence and academic function (Da Fonte & Boesch, 2019; Light & McNaughton, 2014).
The importance of matching an AAC system to the individual has been documented; yet, there is limited research about how professionals select such systems (O’Neill & Wilkinson, 2020; Schlosser & Raghavendra, 2004). Therefore, to effectively implement evidence-based practices in the selection of an AAC system relevant parties should consider current evidence, while also acknowledging the expertise of the professional, and the perspectives of the relevant parties (Light & McNaughton, 2014). The absence of empirical evidence may lead to erroneous decisions that could inhibit the individual’s communication development. As a result, feature matching becomes essential to effectively identify AAC systems. Feature matching involves an ongoing assessment process that identifies the current skill level of the individual and provides suggestions that match the individual’s ability to the features of an AAC system (Light & McNaughton, 2014). To implement this process, the evaluator should have an understanding of the different systems available, the features of the system, and of the individual’s skills. Currently, a web-based interactive tool exist that follows these recommendations, the Student Inventory for Technology Supports (SIFTS; Ohio Center for Autism and Low Incidence, 2012). The SIFTS is a promising tool to identify AAC systems. Although it does not provide users with a specific app to use, it provides the features that are necessary for the user and an application can then be suggested by a service provider.
Mobile devices when used with AAC apps are considered non-dedicated, high technology systems that allow a person to communicate through synthesized or digitized speech. Given the widespread availability, popularity, and the increased trend of such system, the use as an AAC system will most likely continue to increase (Light & McNaughton, 2014). At-present, AAC apps are readily available through various app stores including Apple, Google Play, Galaxy Apps, and Amazon. While there may be many AAC apps, currently, there is no clear process on how to deem if an app is appropriate for a specific individual to use and if it meets the needs of that individual. Currently, most AAC apps have little to no empirical evidence to support their use. To further muddle the situation, relevant parties are relying on company ratings, customers’ reviews, or the description provided by the company on the app (Cherner et al., 2014).
To determine the availability of empirical app evaluation systems, Boesch and colleagues (in review) conducted a literature review to identify studies that had created and empirically evaluated app evaluation tools. They found that there were no evaluation tools specific for AAC apps, but several existed for educational apps. From the 11 studies included in the Boesch et al. (in review), only 4 app evaluation systems were empirically tested (Lee & Kim, 2015; Papadakis et al., 2017; Sanromà-Gimenéz et al., 2021; Weng & Taber-Doughty, 2015). For example, Lee and Kim (2015) developed the Criteria Model for Evaluating Educational Apps, a rating scale to analyze apps based on the variables of teaching and learning, screen design, technology, and economy and ethics. Papadakis et al. (2017) created a rubric, Rubric for the EValuation of Educational Apps for preschool Children (REVEAC) that evaluates preschool educational apps in four areas: contents, design, functionality, and technical quality. The rating scale evaluation tool designed by Sanromà-Gimenéz et al. (2021) could be used for a variety of educational apps and was presented in the form of a questionnaire consisting of 36 items grouped into six sections. Finally, Weng and Taber-Doughty (2015) tested the user friendliness of a prototype rubric designed for screening and evaluating apps intended for educational use. Out of the final 11 articles that met the inclusion criteria set by Boesch et al. (in review), only 3 articles met over half of the methodological rigor criteria (i.e., Papadakis et al., 2017; Sanromà-Giménez et al., 2021; Weng & Taber-Doughty, 2015). But after evaluation, only one article, Weng and Taber-Doughty (2015), had a rigor of over 80%. The lack of rigor present in the rest of the articles reduces their validity for the study.
Despite the lack of empirically validated AAC app evaluation systems, there are several publicly available, non-empirical systems. Gosnell et al. (2011) created the Feature Matching Communication Applications, a rubric designed to help match a person’s strengths and needs to AAC apps. Marfilius and Fonner (2012) created multiple feature matching checklists for different types of AAC systems. One of these checklists, Feature of AAC APPs, focuses on AAC apps, and includes dimensions such as input features, visual appearance, processing features, output features, and scanning features. Similarly, Parker and Zangari (2012) created the Rubric for Evaluating the Language of Apps for AAC (RELAAACs) that uses a five-point Likert scale to evaluate an app’s ability to support 16 communication skills across two domains, functional communication learning and language learning. The iEvaluate App Rubric by Van Houten (2011) attempted to help in identifying effective apps for an individual with complex communication needs based on their individualized education program goals.
While these app evaluation systems may be used and accepted by relevant parties, the lack of empirical AAC app evaluation systems continue to exist, and it is problematic. To attempt to decrease this gap, the purpose of this study was to create an empirically validated AAC app evaluation tool that could be used regardless of the evaluator’s knowledge and skill levels in AAC. This study’s research questions are: (a) how reliable is the AAC app evaluation tool? (b) to what extent are there differences in reliability across app communication skill categories? (c) to what extent are there differences in reliability across the dimensions within the AAC app evaluation tool? (d) what are the perspectives of novice raters on the AAC app evaluation tool?
Method
The goal of this study was to design and pilot an AAC app evaluation tool. The purpose was to offer professionals a comprehensive tool to help with the evaluation and identification of communication apps for individuals with complex communication needs. The reliability of the tool was collected and analyzed across two phases: Phase 1, the pilot phase with expert raters (members of the research team); and Phase 2, the testing phase with both, novice raters and experienced novice raters (novice+). Data were analyzed within and across phases, communication app categories, and dimensions of the AAC app evaluation tool.
Sampling and Recruitment
Institutional Review Board (IRB) approval was obtained prior to the recruitment or dissemination of study materials. Convenience sampling was used to reach accessible participants who fit the inclusion criteria. Participant inclusion criteria encompassed: (a) enrolled in a low incidence training program; (b) seeking initial teaching licensure or teaching endorsement in severe disabilities; and (c) enrolled, at a minimum, in an introductory level course in special education for severe disabilities.
To recruit potential participants, an email was distributed to all teacher candidates pursuing an undergraduate or master’s degree in special education, specifically, teacher candidates from a low incidence training program at Southern University (n = 43). Potential participants within this track were completing a teacher preparation program that specializes in training future special educators with the knowledge and skills needed to teach individuals with severe disabilities, including individuals who present needs in the areas of “general learning, personal and social skills, and/or sensory and physical development” (Westling et al., 2020, p. 4). Follow-up emails were sent every 2 weeks across a month, to those who did not respond to the initial communication, with a final response rate of 69.77%.
Participants
To collect demographic information on the participants, a rater profile was created and disseminated using REDCap™ (Research Electronic Data Capture), a web-based software platform used to manage online surveys and support data collection for research studies (Harris et al., 2009). The rater profile included nine questions related to: (a) participant’s demographic information (n = 5), which consisted of questions on age, gender, ethnicity, degree seeking, and year in enrolled program; (b) practica experiences with individuals with complex communication needs (n = 1), which asked information on participant’s exposure to individuals with complex communication needs through a field experience; (c) coursework in augmentative and alternative communication (AAC; n = 1), which asked if they had previous coursework related to AAC; (d) experience with individuals with complex communication needs outside of any practica (n = 1); and (e) teacher licensure (n = 1), which collected information on whether the participant was seeking initial teaching licensure or a teaching endorsement. If participants indicated they were seeking teaching endorsement, two additional questions were asked: (a) the type of initial teaching licensure they currently held; and (b) if they had years of prior teaching experience.
Participant Demographics.
Note. + = novice+ raters; - = not applicable; A = Asian; AC = augmentative and alternative communication academic course; B = Black; BD = bachelor’s degree; E = seeking teaching endorsement; H/L = Hispanic/Latinx; I = seeking initial teaching licensure; MD = master’s degree; NA = Native American; O = Other; P = practicum or field experience in augmentative and alternative communication; W = White.
Research Design
Usability testing (Barnum, 2020; Rubin & Chisnell, 2008) was used to collect qualitative and quantitative data from common users, in this case, future special education teachers, during a task that may be part of their professional role or responsibilities (Rubin & Chisnell, 2008). The independent variables in this study were the participants’ characteristics and the dependent variables were the validity and reliability of the AAC app evaluation tool. The overarching goal was to establish reliability of expert raters and determine differences among novice and novice+ raters (Bias & Hoffman, 2014).
Application Selection
The CALL Scotland’s (2021) iPad Apps for Complex Communication Support Needs: Augmentative and Alternative Communication was used to identify communication apps. The tool classified the communication app in three categories, including: (a) communication skills, which incorporated apps that support the development of pre-requisite communication skills (see Heimann et al., 2006; Maatta et al., 2011), early vocabulary and language skills, and some with communication skill building approaches (e.g., PECS); (b) simple communication, which comprised apps that aid early generalization of language skills (see Dyches et al., 2002), provide the opportunity for choice-making, environment-specific vocabulary, and the creation of personal or social stories. Simple communication apps contain limited, if any, starter content, and allow the use of photographs, line drawings, and individualized speech-recording messages; and (c) full communication, included apps that are text or symbol based, and provide a wide range of vocabulary and built-in symbol library, and provide a message bar to allow for sentence building. For this study, 18 apps were randomly selected from the CALL Scotland and evaluated by the expert raters. The ratings of the apps were reviewed from Apple App Store® (ranging 1 to 5; M = 4.03), Google Play Store™ (ranging 1–5; M = 3.75), and Jane Farrall Consulting (ranging 1 to 3; M = 2.5). The price of the identified apps ranged from free (n = 7) to paid apps (n = 11), with a cost ranging from $4.99 to $299.99 (M = $110.90). Of the 18 apps, 9 apps were randomly selected with a representation of 3 apps per category. The apps were then divided into three app groups for their evaluation among novice and novice+ raters.
AAC App Evaluation Tool
The AAC app evaluation tool was designed as a rating scale, as the intent was to collect information on the degree to which a list of specified characteristics was evident on each of the selected apps (see Brookhart, 2013). Questions were determined based on a feature matching model in combinations with the dimensions suggested by Boesch et al. (in review). The AAC app evaluation tool consisted of two sections, (a) the app information section, which collected data on the app’s background and application design features; and (b) the app evaluation sections, which was divided into three dimensions, including usability, display, and output. The first draft of the app evaluation tool was comprised of a total of 30 questions, 14 in the usability dimension, 14 in the display dimensions, and 2 in the output dimension. Factor analyses were conducted by the research team (co-authors) to determine common correlates among factors. An expert in statistical analyses was consulted as needed for final decisions. Results suggested one common factor with 77.54% variance, suggesting that all questions were related. Questions were then categorized across three categories (dimensions). Cronbach Alpha’s were calculated for each dimension, ⍺ = .924 for the usability, ⍺ = .962 for display, and ⍺ = .948 for output. The final AAC app evaluation tool consisted of a total of 10 questions, 3 for usability, 5 for display, and 2 for output. Participants completed the AAC app evaluation tool on REDCap™, an online research electronic data capture platform. This platform was also used to collect (a) raters’ demographic information (n = 12); and (b) social validity questionnaire, with 7 Likert scale questions, and 1 open-ended.
Data Analysis
Both formative and summative data were collected throughout the study. Data were analyzed to determine differences between expert, novice, and novice+ raters. Data were collected and analyzed during the development of the tool, and to evaluate the reliability, internal consistency, and validity of the AAC app evaluation tool. Social validity was collected by analyzing participants’ perspectives on the AAC app evaluation tool.
Internal Consistency, Reliability, and Validity
Internal consistency, reliability, and validity were calculated both globally (one construct) and locally based on communication app categories and the dimensions within the AAC app evaluation tool. Using SPSS (version 27.0), Fleiss’ kappa and agreements within ±1 point (inter-rater adjacent agreements) were calculated to assess reliability within and across phases. Fleiss’ kappa was used to determine reliability between raters. Using Cicchetti’s (1994) standards, the criteria set was poor (0.00–0.40), fair (.41–.59), good (.60–.74), and almost excellent (.75–1.00). The percentages of inter-rater adjacent agreements were set as acceptable (between 75% and 89%) to high (≥90%) implementing the standards suggested by Bajpai et al., 2015. To determine internal consistency, Cronbach’s alpha was calculated for the AAC app evaluation tool with a coefficient of variation within .910 (95% confidence interval). Cronbach’s alpha criteria was set as ‘good (0.8 ≤ α ≤ 0.89) to excellent’ (α ≥ 0.9) using the standards of Gliem and Gliem (2003).
Procedural Integrity
A procedural integrity checklist was created to ensure consistency in the information provided to and interactions with the participants. The checklist consisted of four steps, that include: (a) initial email sent with recruitment information; (b) follow-up email to potential participants who did not answer the initial recruitment request; (c) a thank you email for indicating interest to participate and confirming their agreement and time commitment; (d) email to participants with information on the purpose of the study, instructions on tasks to be completed, the rater number, the iPad number, password to access the iPad, a list with the three apps to be evaluated, and the REDCap link where participants would consent their participation, complete the rater profile, and evaluate the apps using the AAC app evaluation tool; (e) post evaluation, an additional email was sent with a REDCap link with the social validity survey Supplemental Figure 1.
Procedures
A top-down and bottom-up method of organization were used to develop, design, and finalize the AAC app evaluation tool (Barnum, 2020). The top-down approach is used to organize findings based on predetermined criteria. Whereas, a bottom-up approach is used to organize the data into groups, categories, or themes (Barnum, 2020). To determine the dimensions within the AAC app evaluation tool, the top-down approach was used. Using the systematic literature review conducted by Boesch et al. (in review), a top-down approach was used to determine five broad dimensions that were evident across most of the identified literature, including: (a) background information; (b) design features; (c) usability; (d) individualization; and (e) overall impressions. An additional search was conducted to identify other available app evaluation systems that did not meet Boesch et al.’s inclusion criteria. Using a bottom-up approach, questions on the AAC app evaluation tool rearranged based on factor analysis results.
A first version of the AAC app evaluation tool was designed including three dimensions outline in the available literature: (a) background (n = 16 questions); (b) usability (n = 14 questions); and (c) individualization (n = 14 questions). All items were rated on a 5-point Likert scale, except for the background information dimension. For consistency across participants, operational definitions were provided for the unipolar labels used in the Likert scale (e.g., Steinberg & Rogers, 2020). This consisted of: not at all = 1, participant did not agree the app aligned with any part of the statement; slightly = 2, participant agreed the app aligned marginally with the statement; somewhat = 3, participant agreed the app aligned with some part of the statement; very much = 4, participant agreed the app aligned with most of the statement; and extensively = 5, participant agreed the app aligned with all of the statement.
Prior to data collection, two in-service special education teachers were asked to review the AAC app evaluation tool to obtain feedback on format and wording. Both teachers earned a master’s degree and held a special education teacher licensure. One teacher had 25+ years and the other had 13 years of teaching experience. Both teachers served individuals with severe disabilities, including individuals with limited functional speech. One of the teachers held the primary role in the identification of potential assistive technology systems for individuals to be able to actively participate and access classroom activities. Based on the reviewers’ suggestions, the AAC app evaluation tool was revised by adding an option for switch access within the background dimension.
Once reliability and internal consistency was reached among expert raters, factor analysis was conducted to determine variability within and across dimensions to guide the potential redistribution of questions within each of the dimensions and the potential relabeling of dimensions. Specifically, using the literature review findings, it was determined that a dimension for the app’s individualization capabilities should be included with an app evaluation. Yet, results from factor analysis indicated that these questions should be placed under the app display and app output dimensions. Therefore, revisions were made to the AAC app evaluation tool in five areas based on reliability, internal consistency, and factor analysis data.
The final AAC app evaluation tool incorporated a total of 23 questions and 5 dimensions. The dimensions included the app background information (n = 12 questions) which collected information using a checklist on the type of communication app and the target skills. The app usability (n = 3 questions) which gathered information on the type of communication functions and the timing of the speech output after selection using a rating scale with a Cronbach of 0.913. The app display features (n = 5 questions), which included questions related to the app’s capabilities for customization, such as the type of symbols, number of cells per page, and text size using a rating scale with a Cronbach of 0.89. The app output (n = 2 questions) which collected information on the degree to which the speech output can be individualized using a rating scale with a Cronbach of 0.906, and the overall impression (n = 1 question), which was the participant’s holistic opinion of the app using a rating scale.
Pilot Phase
In the pilot phase, expert raters evaluated the apps using the AAC app evaluation tool. Expert raters were involved in the design, development, and revision process of the AAC app evaluation tool, and included three white females between the ages of 24–25 who were completing a master’s degree in special education (low incidence disabilities track) and were part of the research team. Two out of the three expert raters had a bachelor’s degree in special education, they all had 2–3 years of teaching experience and had taken a master’s level AAC course.
During the pilot phase, 18 apps were evaluated, across the communication app categories of communication skills (n = 4), simple communication (n = 5), and full communication (n = 9). The goal was to determine which apps were rated as ‘high,’ ‘medium,’ or ‘low’ within each category. The classification was determined based on the average rating score among experts’ raters post using the evaluation tool. For this study, a rating of high, was defined as apps that received scores between 3.67 and 5; medium was defined as apps that obtained scores between 2.68 and 3.66; and a low rated app was referred to as an app that received scores between 1 and 2.67. The purpose was to select communication apps to be evaluated by novice raters that represented a range of rating scores to determine if the AAC app evaluation tool was reliable across different qualities of apps.
Testing Phase
Based on the results from the pilot phase, three sets of communication apps were selected for a total of 9 total apps, including one app per communication app category and one app per quality rating (3 apps per set). In the testing phase, there were two participants groups. Novice raters included participants with no prior exposure to the AAC app evaluation tool and evaluated a total of three apps. Novice + raters were a sub-group of novice raters who were asked to evaluate the additional sets of apps (6 additional apps).
Novice Raters
Novice raters were divided into three groups with 10 participants per group. They were randomly assigned to a set of apps. Each participant was given 1 week to complete the evaluation of the three communication apps.
Novice+ Raters
To further test the reliability of the AAC app evaluation tool, 12 of the novice raters were randomly selected to evaluate additional apps (novice+). The purpose was to determine if having additional practice using the AAC app evaluation tool would influence the reliability of the novice+ raters. Novice+ raters were contacted via email to determine their interest in participating further. All randomly selected novice+ raters agreed to evaluate six additional apps. To ensure procedural fidelity, each participant received an email with the rater number, the iPad number, password to access the iPad, a list with the six additional apps to be evaluated, and the REDCap link to evaluate the apps using the AAC app evaluation tool.
Social Validity
Social validity was collected through a survey designed to gather information on participants’ perspectives about the communication app evaluation system. The survey was disseminated using REDCap™ after the completion of the app evaluation. To validate the survey prior to its dissemination, the first version was sent to one reviewer, who was asked for feedback and authentication of the questions, and pilot test the survey (Robinson & Leonard, 2019). The reviewer was a university faculty professor with a doctoral degree in developmental psychology, with extensive knowledge and experience in survey research, and who was not directly related to the study. Based on the feedback, revisions were made to increase clarity of two questions. Two members of the research team piloted the survey to estimate completion time (5–10 minutes; data were not included for analysis).
The final version of the survey consisted of two sections (top-down), the user’s experience (n = 7 questions), and recommendations (n = 2 questions). The user’s experience included questions related to the participants’ overall perspectives on the communication app evaluation system. In this section, participants responded using a 5-point Likert scale to the extent to which they agreed with the statements about the communication app evaluation. The recommendations section included an open-ended question that asked the participant to identify any potential changes that could be made to the AAC app evaluation tool.
Triangulation analysis was used to assess multiple perspectives of the data collected in relation to social validity (Barnum, 2020). The recommendations of novice and novice+ raters were analyzed based on common themes, using a bottom-up approach with a ‘perfect’ agreement (k = 1). One-way ANOVAs were used to calculate descriptive statistics for each item on social validity survey. Independent t-tests were utilized to determine significance between summed user experience scores and demographic characteristics, and repeated measures ANOVA were used to compare questions to each other to find if there was any significance between questions.
Results
Pilot Phase Data
Reliability and Internal Consistency of the Pilot and Testing Phase.
Testing Phase Data
Global findings for novice raters indicated an inter-rater adjacent agreement below the acceptable criteria (70.44%). The reliability was ‘poor’ (k = .23); however, the internal consistency was ‘excellent’ (⍺ = .909). Locally, the inter-rater adjacent agreement was also below the acceptable criteria for all dimensions. The reliability among novice raters was ‘poor’ across all dimensions. Yet, internal consistency was ‘good’ in the display dimension, and ‘excellent’ in the usability and output dimensions. When viewing results on the app categories, the full communication app category had ‘acceptable’ inter-rater adjacent agreement, but the communication skills and simple communication app categories were below the acceptable criteria. The reliability was ‘poor’ across all three app categories; still, internal consistency was ‘good’ across app category (see Table 2).
Global findings for novice
+
raters indicated that inter-rater adjacent agreement was below the set criteria (71.39%). Reliability across participants was ‘poor’ (k = .187), but the internal consistency was ‘excellent’ (
Social Validity Data
All participants (n = 30) responded to the multiple-choice section of the social validity survey. In this section, participants ‘agreed’ that the evaluation system could be completed in a reasonable amount of time (M = 4.30; SD = 0.596), that the app evaluation tool was easy to use (M = 3.67; SD = 0.994), and that it was useful (M = 3.67; SD = 0.711). Results suggest that there was a difference between participants who had taken an AAC course prior to using the AAC app evaluation tool, and those who had not, t(30) = 0.597, p = 0.021. Findings also suggest that participants who had taken an AAC course had a higher overall satisfaction (M = 25.60; SD = 3.89) with the AAC app evaluation tool than participants who had not taken an AAC course (M = 24.40; SD = 6.74). Only 60% of the participants responded to the open-ended question of the survey. The most suggested recommendation (77.78%) was to include operational definitions to help clarify the terminology used in the AAC app evaluation tool. Yet, there was no significant relationship between participants’ backgrounds and the recommendations provided.
Discussion
Although several app evaluation tools have been created and evaluated (Lee & Kim, 2015; Papadakis et al., 2017; Sanromà-Gimenéz et al., 2021; Weng & Taber-Doughty, 2015), up until this point, all app evaluation tools specific to AAC were non-empirically based despite being widely used (e.g., Gosnell et al., 2011; Marfilius and Fonner, 2012; Parker and Zangari, 2012). This study is the first to incorporate dimensions from other empirically-test tools (i.e., usability, display, and output) and test the reliability of an AAC app evaluation tool formatted as a rating scale. Specifically, this study aimed to assess the reliability of an AAC app evaluation tool across two phases: a pilot phase, and a testing phase. Expert raters used the tool during the pilot phase while novice raters and experienced novice raters (novice+) used it during the testing phase to ascertain its reliability.
The team measured reliability using inter-rater adjacent, Cohen’s kappa, and internal consistency. Pilot testing data suggested that expert raters met criteria, indicating that the AAC app evaluation tool had ‘acceptable’ to ‘high’ agreement. Yet, within the novice rater group (i.e., novice vs. novice+) the majority of comparisons showed ‘below acceptable’ agreement for the app types (i.e., full communication, simple communication, communication skills), and dimensions (i.e., usability, output, display). Given that expert raters demonstrated greater reliability measures than novice raters, it is likely that this difference was due to several factors. First, the expert raters were involved in creating the AAC app evaluation tool. Thus, they had more exposure to the tool, were keenly aware of what each component of the tool was designed to evaluate and were knowledgeable of the terms used in the tool. Expert raters also evaluated the greatest number of AAC apps, which further allowed them to spend the most time using the tool as compared to the novice raters. These findings are aligned to the assertions by Gronlund and James (2005) indicating that practice increases reliability. Similarly, Semmelroth and Johnson (2014) conducted a study in which they tested the reliability of a teaching evaluation tool. They found that multiple opportunities to use the tool was an important factor in ensuring acceptable levels of reliability.
When comparing the novice raters’ reliability across all app types and dimensions, kappa scores were ‘poor’. Interrater adjacent reliability was ‘below acceptable’ for the novice and novice+ raters with scores in the ‘acceptable’ range for only a few. Collectively, these findings indicate that novice raters were not reliable in using the AAC app evaluation tool. Similar to a study by Semmelroth and Johnson (2014), the current findings from expert and novice raters suggest that reliability increases with more practice and opportunities to use the tool. Closer inspection of the reliability data also pointed to differences between app categories (i.e., simple communication, communication skills, and full communication). Novice raters were more reliable when evaluating apps categorized as full communication apps when compared to apps categorized as communication skills or simple communication. It is possible that novice raters achieved higher reliability with full communication apps due to the greater opportunities to individualize such apps, and therefore greater number of app features to evaluate. For example, for apps classified as ‘communication skills’ type, most of these apps served one function (e.g., requesting). As a result, there were fewer opportunities to individualize the app features such as its symbol size, text, vocal output when compared to ‘full communication’ apps. Thus, raters were more reliable in apps that had more customizable functions (Abdelmoula et al., 2015).
Given that special education teachers and other professionals may have a profile similar to the novice raters in this study, it is important to identify ways to improve the reliability when using this app evaluation tool. One strategy could be to provide more training on the AAC app evaluation tool. Another strategy could be to provide users with operational definitions to provide further guidance on how to use the tool and potentially increase knowledge of the tool to be comparable to that of the expert raters. Based on the social validity data obtained for the current study, novice raters indicated that having operational definitions for AAC-specific terms would have been helpful in using the AAC app evaluation tool. This finding aligns to that of Gronlund and James (2005) in which they found that in order for raters to become more reliable, an understanding of the language with operational definitions was necessary. Nonetheless, despite the lack of operational definitions, novice raters in the present study indicated that overall, they enjoyed using the tool and found it useful.
Another important finding pertained to the tool’s internal consistency. Because many professionals are often limited in time due to large caseloads or high number of students they work with (Russ et al., 2001), it is important to ensure that the tools they are using have the appropriate number of dimensions and relevant items. Therefore, the internal consistency of the current AAC app evaluation tool was tested, and data were generally found to range from good to excellent. These findings indicate that the items on the app evaluation tool contained a sufficient number of relevant items, thereby decreasing the user’s time needed to spend completing sections that are redundant or unnecessary. This is important to note given that other evaluation tools commonly used by professionals have not been assessed for internal consistency.
Practical implications
Based on the study’s findings indicating that indirect training among team raters likely led to higher reliability than novice raters, it is recommended that professionals who use the AAC app evaluation tool familiarize themselves with this tool. Although it is unclear the type and depth of training needed to increase the reliability among novice raters, professionals may require training and practice prior to assessing AAC apps for their students or clients. Otherwise, the findings of the AAC app evaluation tool should be used with caution if no training is obtained. Results also indicated that raters wanted operational definitions included in the AAC app evaluation tool. Therefore, it is recommended that as part of familiarizing themselves with this tool, professionals should learn key terminology to ensure they have the content knowledge needed to successfully evaluate AAC apps.
Limitations and Future Research Directions
Although findings may be promising, replications are necessary as there were three overarching limitations to this study. First, the sample of this study consisted of pre-service teachers’ candidates. While pre-service teachers’ candidates are an essential group of participants, as they will be involved in the decision-making process, a diversified and larger sample should be considered. Future research could consider evaluating differences and similarities in the evaluation of apps among pre- and in-service professionals. For example, future research could compare if there are differences among special education teachers and speech language pathologists, as they are critical in the service provision of individuals with complex communication needs. Likewise, future research could also compare the reliability among professionals, families, and users, when using an AAC app evaluation app tool. Outcomes of such research may produce essential information in the app evaluation and feature matching process.
Second, while the AAC app evaluation tool was comprehensive and easy to use, raters did not receive training on the use of the evaluation tool. Although findings suggest that the AAC app evaluation tool could be used without any training, training may have been a component that impacted the reliability among expert and novice raters. A potential area of future research could be to evaluate if there are differences across a variety of training methods, such as provision of operational definitions (written training), versus video-training, versus in-person training (professional development), on raters’ reliability. A comparison across a variety of training methods could help determine the effectiveness of the tool, while at the same time evaluating the level of knowledge needed to use such tools. Third, during this study, the AAC apps were evaluated without a particular user in mind; raters were simply asked to evaluate the apps. As such, the possibility exists that raters may be able to evaluate apps differently, if they had a user in mind, specifically when evaluating the dimensions of usability and individualization. Therefore, the possibility exists that the evaluation of apps with for a specific user, may in turn enhance reliable. As result, future research is needed to determine if reliability is influenced when evaluating an app for a target user versus just the app capabilities.
Overall, the need for a valid and reliable app evaluation tool is essential, especially given that AAC apps will continue to grow in popularity. Therefore, further research must be conducted to create reliable app evaluation tools that professionals of varying knowledge and skill levels can utilize with high reliability and accuracy. Without empirically evaluated app evaluation tools, it is challenging for professionals and families of individuals with complex communication needs to identify the most suitable and effective communication app.
Conclusion
The aim of this study was to create an AAC app evaluation tool and evaluate its reliability and social validity. The goal was to determine if this tool could be used across AAC apps and to determine if reliability differed among expert and novice raters. Results suggest that novice raters were not as reliable as experts, and therefore, the possibility exist that with additional practice or training, novice raters may become more reliable. Unfortunately, the lack of empirically validated app evaluation tools proves a challenge to practitioners and relevant parties in determining what apps will be the most appropriate for individuals with complex communication needs. The lack of validated app evaluation tools can increase the possibility of system abandonment. As such apps should be evaluated prior to its use to ensure the app is selected based on the individual’s unique abilities and needs.
Supplemental Material
Supplemental Material - Designing A Valid and Reliable AAC App Evaluation Tool: Differences Between Team and Novice Raters
Supplemental Material for Designing A Valid and Reliable AAC App Evaluation Tool: Differences Between Team and Novice Raters by Miriam C. Boesch, M. Alexandra Da Fonte, Melissa J. Cavagnini, Kaitlyn R. Shaw, Keren E. Deneny, and Margaret F. Davis in Journal of Special Education Technology
Footnotes
Acknowledgments
We would like to give a special thanks to Stephanie Camacho, Gillian Neff, and Jennifer Lipof for their support and feedback during the editing phases. We would like to also thank Emily DeLuca for her support with the data analysis and during the early phases of this project. A special thanks also goes out to Nicole Wolfe and Emily Sheridan for their contributions during the initial stages of this project. Also, we would like to give a heartfelt thanks to all the companies who kindly gave us free access to the communication apps we used in this study.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Supplemental Material
Supplemental material for this article is available online.
Author Biographies
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
