Abstract
Data literacy is essential in today’s workforce, yet existing data literacy assessments lack brevity, limiting their practical use. This study develops a brief data literacy instrument (BDLI) by streamlining the 25-item Global Data Literacy Benchmark (GDLB). Two studies were conducted. The first study (N = 408) explored item reduction and factor consistency using student participants, while the second study (N = 388) replicated findings using a diverse professional sample. Results indicated that a 12-item BDLI achieved similar factor structure and reliability compared to the original GDLB. Confirmatory factor analysis and item response theory confirmed the BDLI’s consistency and discriminatory power. These results suggest that the BDLI offers a practical tool for assessing an individual’s data literacy skills, which can be used to develop targeted learning opportunities at the individual or organizational level. The value and limitations of the BDLI are discussed.
Introduction
Data literacy (DL) is necessary for people to use and consume data intelligently and also enhances people’s relevance-marketability in the workforce (Brynjolfsson and Mitchell, 2017; Schüller, 2022). DL also improves people’s ability to recognize deceptive information and statements (Bond et al., 2006; George et al., 2013), especially in cases where the information is data-rich and data-derived. This, in turn, enables individuals to hold organizations and governments accountable (DATA-REVOLUTION-GROUP, 2014) for false and deceptive information. Consequently, DL allows individuals to become more informed and, thus more fully participate in society (Letouze, 2016; Unesco, 2019). In organizational settings, data-literate employees are more likely to use data-informed decisions that improve organizational processes, enhance decision quality, and lead to identifying new opportunities (George et al., 2013; McKinsey, 2009). This is what is referred to as a “data-driven culture” (Gupta and George, 2016; Mcafee et al., 2012; Ross et al., 2013).
An initial step in striving for such a culture entails first assessing the extent to which organizational members (upper-level executives, middle-level managers, lower-level workers) make decisions based on the insights gleaned from data (Gupta and George, 2016). Yet, many organizations suffer from a shortage of data literate employees (Manyika et al., 2011) and lack the expertise or assessment tools to evaluate the data literacy of their organization’s members (Cui et al., 2023). Whereas consultants and executive coaches are well-equipped to guide these leaders through the data literacy assessment process, there is a dearth of brief data literacy assessment tools.
The 25-item global data literacy benchmark survey instrument (GDLB) is often used to measure DL (DATATOTHEPEOPLE.ORG, 2018). However, its length may make it unsuitable for inclusion in assessments comprising multiple measures or questionnaires. In these instances, instrument brevity is a commodity. Previous studies indicate that research participants are more willing to complete instruments with fewer items (e.g. Carver, 1997; Robins et al., 2002). Additionally, brief instruments are useful when time is limited or there is risk of losing participants’ interest (e.g. Carver, 1997; Robins et al., 2001; Robinson et al., 2003). Crucially, there is evidence that many brief versions of instruments have nearly the same degree of reliability and validity compared to their long-form versions (e.g. Burisch, 1984, 1997).
This project had two objectives: (1) reduce the number of items required to reproduce the GDLB’s results and (2) explore whether samples of participants across different populations provide similar responses to this briefer version of the GDLB. These objectives were accomplished across two studies. Study 1 explored whether a reduced set of items derived from the GDLB would load on similar factors using a student sample. Study 2 replicated the first study’s results with a large, diverse sample of professionals whose jobs require some degree of data literacy. Lastly, we examined 15-item and 12-item versions of the instrument with a composite of the student and professional samples from Study 1 and Study 2.
Literature review
This literature review provides a comprehensive overview of data literacy, highlighting its importance, challenges, and potential strategies for improvement. By addressing these aspects, stakeholders can better understand how to cultivate data literacy in various contexts.
Importance of data
Describing data as the “new oil” speaks to the newfound importance of data literacy in the 21st century for individuals and society (Holeni, 2020). Stepping back and looking at basic literacy, one finds that many across the world do not possess basic literacy skills, limiting their ability to fully engage and participate in society (Letouze, 2016; Unesco, 2019). The lack of data literacy skills challenges individuals and organizations, impacting their ability to successfully compete in today’s data-centric world. Yet, when did data become important?
The collection and analysis of data represent core steps in a research methodology, supporting discovery for hundreds of years; data predates computers. Researchers recognized the importance of information literacy and computer literacy in the 1970s and 1980s (Noble, 1984; Spitzer et al., 1998), yet the focus on data literacy did not appear until after the turn of the 21st century (Hunt, 2005; Schield, 2005; Stephenson and Schifter Caravello, 2007). Why did this take so long?
The focus on data literacy coincides with the convergence of several key technical capabilities. Continued increases in computer processing power, the ability to store and process vast amounts of data, and the ability to leverage advanced algorithms did not become widely available until around 2007 (Friedman, 2017). And while the conceptual underpinnings of advanced analytics, such as artificial intelligence, trace back to the 1940s (Buchanan, 2005), after this convergence, organizations moved quickly to leverage data and analytics for competitive advantage (Davenport, 2006; Lavalle et al., 2011; Manyika et al., 2011). Organizations scrambled to leverage the new capabilities associated with these advancements.
To manage the rapid growth of data and analytics and develop their data resources as a strategic asset, organizations created the Chief Data Officer (CDO) position (Xu et al., 2016). From the first recognized CDO in 2003 (Zhang et al., 2017), 90% of large organizations now employ a CDO (Gartner, 2016). The US government also recognizes the potential value available from internal data sources, as evidenced by a 2019 law mandating that federal agencies adopt the CDO role (Shibu, 2019).
The widespread adoption of the CDO role presents a unique view of the importance of data versus other major technology initiatives, such as Enterprise Resource Planning (ERP), e-commerce, and the Internet. Each of these technological advances greatly enhanced business capabilities, yet the technology was integrated into the organization and did not result in the creation of persistent C-Suite roles like the CDO (Kirby, 2020).
Further, the importance of data continues to dominate organizational priorities. According to recent Society for Information Management surveys, data-related investments, such as AI, analytics, and data integration, remain top priorities for organizations (Johnson et al., 2023, 2024). Enhancing data literacy skills remains a critical priority for organizations.
Challenges associated data literacy
To improve data literacy, one must contend with several challenges: organizational issues, lack of a clear definition, and a lack of existing data-literate individuals.
Many organizations struggle with data-related initiatives. The vast majority of large-scale technology investments fail to achieve desired results, with a 50%–70% failure rate (Deloitte, 2015; Tabrizi et al., 2019; Waid, 2019). For data initiatives, CDOs report similar failure rates (Bennett, 2016). The rapid increase in CDO appointments suggests companies have taken a “knee jerk” approach to harnessing value from data, rather than taking the strategic approach, recognizing that success is less about the technology and more about strategy, alignment, and creating capabilities within the workforce (Davenport, 2006; Feeny and Willcocks, 1998; Pothier and Condon, 2020; Ross et al., 1996). The exemplars that find success with these initiatives do so with a focus on developing digital strategies aligned with their core business strategies, developing a data-centric culture, and creating a data-literate workforce, leaving their less capable peers struggling to find success (Bean, 2018; Dykes, 2019; McKinsey Analytics, 2019).
The lack of a common definition for data literacy presents another critical challenge. Corrall (2019) provided a model highlighting the interdisciplinary nature of data literacy, presenting a complex view of how various forms a literacy intercept with data literacy at the center. The literature suggests that one who is data literate needs to leverage advanced algorithms (Alpaydin, 2016), be curious (Dykes, 2019; Markham, 2020), possess statistical knowledge (Schield, 2005), understand data privacy/security and ethical issues (Bhargava et al., 2015; Markham, 2020), transform raw data into meaning analytical content (Gummer and Mandinach, 2015), understand data quality (Lawson and Desroches, 2019), and understand the digital data infrastructure (Gray et al., 2018), and possess many more skills.
Also, the skills needed to be data-literate vary depending on a person’s role. The skills needed for a data-literate researcher (Koltay, 2017), will differ from those needed for an accountant (AICPA, 2021; Appelbaum et al., 2021), for a journalist (Gray et al., 2012), for business students (Pothier and Condon, 2020), for statisticians (Schield, 2006), and for public servants (Bonikowska et al., 2019), as a few examples. The definitions and skill requirements describe a person with vast specialized skills, or a unicorn. A simpler approach is needed.
Recognizing the confusion surrounding data and analytics, Kirby (2021) attempted to demystify the subject by presenting a framework placing data literacy in context of numeracy/statistics, problems-solving, and domain knowledge, aimed at addressing a specific business issue. The framework further encourages the selective use of the term “data analytics,” by proposing a scaffolding of data skills with the following levels: Basic Data Skills, Data Analysis, Data Analytics, and Data Science.
A skills gap exists. While no shortage exists of data analytics courses (Almgerbi et al., 2022), the skills gap suggests a greater need for basic data-savvy employees rather than high-end data scientists (Manyika et al., 2011). Davenport and Harris (2017) believe that while digital initiatives must start at the top of the organization, a lack of savvy executives persists (Weill et al., 2021), and the gap between firms capable of leveraging data and digital initiatives and those that can not continues to grow (D'ignazio, 2017; McKinsey Analytics, 2019). Most organizations recognize the disruptive nature of data and analytics, but struggle to respond (Downes and Nunes, 2013).
Improving data literacy
Ridsdale et al. (2015) provide a starting point with their comprehensive synthesis of data literacy strategies and best practices, which recognizes the interdisciplinary nature, providing a conceptual framework for the collection, management, evaluation, and application of data-related tasks. This research defined data literacy as “the ability to collect, manage, evaluate, and apply data in a critical manner,” (Ridsdale et al., 2015) which provides a generic definition to cover the disparate approaches found in existing data literacy educational efforts.
The lack of clarity regarding what it means to be data literate remains. College graduates believe they possess data literacy skills (Pothier and Condon, 2020), while employers continue to report that recent college graduates lack basic data literacy skills (Strauss, 2016). A means of assessing data literacy skills will help.
Consistent with the Ridsdale et al. (2015) framework, the “Global Data Literacy Benchmark” (GDLB) (Crofts, 2018) helps organizations of all types assess their level of data literacy, allowing for a focused approach to a starting. Nath and Kirby (2022) prepared factor analysis on this instrument, identifying three factors: (1) Data Analysis, Storytelling, & Decision-Making, (2) Data Acquisition, Navigation, & Quality Assessment, and (3) Data Wrangling, which involves the modeling and enrichment of data.
Data literacy represents a critical skill to actively participate in today’s data-driven world. It enables individuals and organizations to make informed decisions, drive innovation, and improve efficiency. While there are challenges in improving data literacy, various strategies can be employed to overcome these obstacles. By prioritizing data literacy, educational institutions, organizations, and individuals can better harness the power of data to achieve their goals.
Project method
The aim of this project was to create a briefer and more consistent alternative to the GDLB instrument. To accomplish this aim, we sought to develop a brief data literacy instrument (BDLI) by reducing the number of instrument items, improving item parsimony, and improving item clarity.
Reducing instrument items
Nath and Kirby’s (2022) exploratory analysis of the 25-item GDLB identified three factors: Factor 1 (F1) involves data analysis, storytelling, and decision-making skills; Factor 2 (F2) involves data acquisition, navigation, and quality assessment; Factor 3 (F3) involves data manipulation, modeling and enrichment (“data wrangling”). We sought to reduce the GDLB to 15 items or fewer, using an equal number of items for each factor. Our reasoning for reducing the scale to 15 items was based on previous psychometric evidence in scale reduction and pragmatism. This number of items was chosen because brief instruments comprising two, three, and five-item factors can produce similar reliability and validity as their more numerous item counterparts (e.g. Graham et al., 2011; Muck et al., 2007). In some cases, instruments can be reduced to a few items or a single item (e.g. Benet-Martínez et al., 2002; Campbell et al., 1976) without significant losses in reliability. Pragmatically, five items for each factor allows for some assurance in the similarity between factors for weighting purposes, similarity in completion time requirements, and has intuitive appeal as each factor can be used individually across similar study procedures. Items were selected based on the largest within-factor zero-order correlations from Nath and Kirby’s (2022) exploratory factor analyses.
Improving item parsimony and improving item clarity
The principle of parsimony indicates that it is good psychometric practice to connect a single behavior with a single psychological construct. Some of the items in the original GDLB required participants to rate two behaviors simultaneously. This approach may create ambiguity regarding what the participants are rating. For instance, if a participant were asked to affirm the statement, “I am good at interpreting and cleaning data.” In this question, is the participant evaluating their ability to understand data, arrange data in an interpretable way, or both? If both, are these skills equally weighted in the participants’ rating? Thus, we examined the reduced list of GDLB items that described two or more behaviors and then created a new item or multiple new items that attempted to maintain the conceptual core(s) of the original item(s). Lastly, many psychological constructs are difficult to quantify (e.g. uncanniness; Windsor, 2019). Once operational definitions are established, the language used in the psychometric instrument should clearly connect to the operational definition. Thus, psychometric measures must provide users with clear terms and simple language (Simms, 2008). For this reason, we examined each GDLB’s item syntax and revised where necessary.
The instrument uses a capability approach, whereby a participant rates their ability to perform a given task with questions starting with “I can. . .,” followed by a specific data-related task. A seven-point Likert-type scale measures each response from a 1 (Strongly disagree) to 7 (Strongly agree).
Study 1—Student sample
Participants were recruited across three U.S. universities in 2022. The final student sample comprised 408 students. Of these participants, 57% were women, with 44% in the 19–24 age group, 46% in the 25–44 age group, and 10% aged 45 or older. Approximately 20% of the students reported being in year 5+ of school, some of which came from graduate degree programs. The student sample reported a diverse set of majors, including accounting, finance, marketing, business administration, psychology, biology, clinical counseling, cybersecurity, computer science, economics, and engineering. When asked about the quantitative nature of their field of study, 53% indicated either generally non-quantitative or minimally quantitative, while 47% indicated moderately, highly, or extremely quantitative. When asked about their level of comfort with formal analytics, 33% indicated they were either very uncomfortable or somewhat uncomfortable, 58% indicated they were somewhat comfortable, and 9% indicated they were very comfortable.
Study 2—Workforce sample
Working professional participants were invited to complete a Qualtrics survey shared on Linked-In, Facebook, a post to an academic list-serve, and an invitation to members attending a professional association conference; all surveys were completed in the year 2022. An attention check was added to the survey to ensure diligent effort by respondents. The final sample comprised 388 working professionals. This sample reported a diverse set of occupations, ranging from professors, business executives, accountants, software engineers, commercial realtors, school administrators, project managers, counselors, and some 50+ additional occupations. When asked about the quantitative nature of their work, 28% indicated either generally non-quantitative or minimally quantitative, while 78% indicated moderately, highly, or extremely quantitative.
The working professional sample comprised 49% males and 51% females, but generally older, with only 4% in the 19–24 age group, 66% in the 25–44 age group, and 30% age 45 or older. The workforce sample reported higher levels of comfort with formal analytics, with 26% responding that they were either very uncomfortable or somewhat uncomfortable, 52% indicating they were somewhat comfortable, and 22% indicating they were very comfortable.
Study 1 & 2 results
To test the brief data literacy instrument (BDLI) using the factors as identified by Nath and Kirby (2022), a three-factor a priori exploratory factor analysis was conducted using the 15-question instrument. The factor analysis was prepared with the R statistical computing environment (R-CORE-TEAM, 2023) using varimax rotation and retaining factor loadings exceeding 0.40.
The results indicate that the 15-question instrument reported three factors with a total variance of 67.5%, which compares to the original 25-question instrument’s total variance of 63.7%. The Kaiser-Meyer-Olkin (KMO) measure of sampling adequacy value of 0.95 indicates the sample size for this study is adequate (Shrestha, 2021). The results conceptually replicated the 25-question GDLB’s factor loadings.
Table 1 shows the factor loadings for Study 1 and Study 2, with factor loadings for all participants and partitioned subsets for student and workforce respondents. The loadings indicate expected results, with the exception of questions 5, 8, and 15. Question 5 loads to F1 as expected in Study 2 and total results but loads to F2 in Study 1, both with weak correlations. Question 8 unexpectedly loads to F3 across Study 1 and Study 2. Question 15 loads to F3 as expected, yet with weak correlations. These results suggest that Questions 5, 8, and 15 may require removal, creating a 12-question instrument with four questions per factor.
15 Item-version brief data literacy survey factor loadings for Study 1, Study 2..
F1—data analysis, storytelling, and decision-making, F2—data acquisition, navigation, and quality assessment, F3—data wrangling. Factors r < 0.40 not shown.
Composite 12-question results
Next, a composite study was prepared to examine how well the 12-question BDLI loads against the three factors as compared to the 15-question instrument, using the composite sample of 796 responses (408 students and 388 workforce participants), using a confirmatory factor analysis method to expand upon the results of Study 1 and 2.
Factor analysis for the briefer 12-question instrument resulted in three factors with a total variance of 72.9%, which compared to the original 25-question instrument’s total variance of 63.7%. The Kaiser-Meyer-Olkin (KMO) measure of sampling adequacy value of 0.93 indicates the sample size for this study is adequate (Shrestha, 2021).
Table 2 reports the factor loadings for the composite study, with factor loadings for all participants and partitioned subsets for student and workforce respondents. Across the full sample and for the student and workforce subsets, the 12-question instrument produced loadings as expected with reasonable correlations to each factor (r > for all loadings). The superior factor loadings and parsimony of the 12-question instrument support its use. Table 3 reports the mean scores and standard deviation for each item and composite scores for all participants and student and workforce subsets. In the ANCOVA, after controlling for student versus workforce, results showed no significant difference between F1 and F2 (p > 0.05) while showing significant differences between the scores for F1 and F3 (p < 0.001) and between scores for F2 and F3 (p < 0.001). No difference was found between student and workforce scores (p > 0.05). Consistent with Nath and Kirby (2022), across all participants, the highest scores were reported for F2, followed by F1 and F3, respectively.
Brief data literacy survey 12 item-version factor loadings.
F1—data analysis, storytelling, and decision-making, F2—data acquisition, navigation, and quality assessment, F3—data wrangling. Factors r < 0.40 not shown.
Brief data literacy survey results.
CFA and IRT data analysis plan
In Study 1 and Study 2, exploratory factor analysis was used to replicate the three factors found in Nath and Kirby (2022) using our revised items with two unique samples. To examine the instrument using a more rigorous set of tests, the brief data literacy instrument (BDLI) was evaluated using an integrated confirmatory factor analysis (CFA) and item response theory (IRT) approach. Decision-making about whether to further reduce the total number of items in each factor was determined based on CFA model fit indices. If item reduction improved the overall model fit of each factor, the reduced model was kept. The decision to remove each item was based on both Study 1 and Study 2’s EFA results and the CFA standardized error for each item within each factor. For instance, item Q05 was removed from factor 1 because it showed the weakest correlation with the 5-item model in the EFA analysis in Study 1 and Study 2 and the largest standardized error in the CFA (R code is available upon request). Once the factors achieved satisfactory model fit indices, a CFA analysis was conducted on all three factors to determine the total instrument model fit, allowing for correlated factors. Finally, an IRT was performed to replicate CFA results and to determine item spread.
CFA analysis
Lin (2022) recommendations for CFA in R (R-CORE-TEAM, 2023) using the lavaan R package (Rosseel, 2012) were used to examine the model fit indices of the three BDLI factors. Following the Bean and Bowen (2021) approach with insights from Lin (2022), we examined four commonly used fit measures: model chi-square, comparative fit index (CFI), Tucker-Lewis index (TLI), and the root-mean-square error of approximation (RMSEA). The model chi-square test is a significance test used to assess exact sample model fit to the population model. However, it is often considered to be over-sensitive to sample size (Bean and Bowen, 2021; Lin, 2022), so approximate fit indices are often used in addition to estimate model fit. Standards in model fit cutoffs outlined in West et al. (2012) were used. The comparative fit index (CFI) and Tucker-Lewis Index (TLI) are used to evaluate approximate model fit using an incremental approach, with a standard cutoff of 0.95 or higher. The root-mean-square error of approximation (RMSEA) is an approximate model fit index that uses an absolute approach. The RMSEA point estimate for close fit is ⩽.05, and reasonable fit is ⩽.08, with a 90% confidence interval range of 0.01–0.08 (Lin, 2022).
Based on the EFA results of Study 1 and Study 2, each five-item and four-item version of the three BDLI factors were analyzed. According to the model chi-square results, Factors 1 and 2 did not show exact model fit for either the 5-item or 4-item model (p < 0.05). According to the model chi-square results, Factor 3 showed exact model fit (p > 0.05) for both the 5-item and 4-item models. All models showed adequate CFI and TLI scores (>0.95). The 4-item models for Factors 2 and 3 showed reasonable approximate fit (Lin, 2022), but the upper confidence intervals exceeded recommended cutoffs. Factor 3, for both the 5-item and 4-item models, met all recommended cutoffs. However, the 4-item model did show improvement for each fit index, which makes it preferred to the 5-item model.
The three-factor 4-item model of the BDLI did not show exact model fit but did show adequate CFI and TLI values (>0.95), and the RMSEA value showed reasonable approximate model fit (0.056) as well as acceptable 90% confidence intervals. These results led us to use IRT analyses on each factor’s 4-item version. See Table 4 for details.
CFA model fit indices for the three factors.
IRT data analysis
Following guidelines from Bean and Bowen (2021), we conducted IRT analyses for each factor using the R statistical computing environment (R-CORE-TEAM, 2023). We assessed item fit for each of the factors using the mirt (Chalmers, 2012) and GAIPE (Bean and Bowen, 2021) R packages. Using the index S-χ2, we examined IRT slope and location parameters, and visually inspected the IRT plots, as recommended by Bean (2021) and Kılıç et al. (2023).
Table 5 provides the item graded response model parameter estimates for items within the three factors. The slope parameter (a) provides a measure of how well each question discriminates the measured level of each factor trait. Using the threshold of a > 1.75 as suggested by Baker (2001), the results indicate the questions within each factor provide a very high level of discrimination. Q02 (Analyze data for decision-making) produces the greatest discrimination for Factor 1, while Q09 (Access multiple types of data) and Q13 (Correct errors within data sources) produce the greatest discrimination for Factors 2 and 3, respectively. Thus, these overall results provide good evidence that the BDLI is a potential tool for assessing an individual’s DL strengths and weaknesses. Appendix A contains the 12-item BDLI.
Item graded response model parameter estimates.
The (b) parameters in Table 5 report the theta level at which an individual has a 50% likelihood of selecting the given point on the Likert scale. Figures 1–3 provide the item probability function plots for each factor. These plots demonstrate the steady progression, by level ability, for each item.

Factor 1 item probability estimates.

Factor 2 item probability estimates.

Factor 3 item probability estimates.
Discussion
The present study aimed to test the consistency of the factors identified by Nath and Kirby (2022), using a shortened version of the GDBL. The factor analysis results across the diverse student and workforce samples demonstrate the briefer instrument produces the same underlying factors found in the GDBL. These findings were consistent across student and workforce samples, using the 15-question and 12-question instruments, with better results from the 12-question instrument.
The three factors of F1—Data Analysis, Storytelling, and Decision-Making, F2—Data Acquisition, Navigation, and Quality Assessment, and F3—Data Wrangling identify the work needed to leverage data for greater insight and decision-making. F1 involves applying the appropriate analytical techniques, algorithms, and models, and creating compelling messages to support the analysis and decision-making process. F2 involves the sourcing, collection, and assessment of the veracity of the data, which is often stored in multiple locations, systems, and formats. F3 involves cleaning and modeling processes necessary to support the analysis and decision-making processes. These factors reflect the core capabilities needed for someone that is data literate. Furthermore, the BDLI provides a briefer and more rigorously validated instrument. This instrument could be used to explore both basic research questions regarding data literacy and assessment questions regarding individuals’ data literacy. Our results support use of the 12-item BDLI by academics and corporate trainers to measure data literacy.
As organizations pursue increased data literacy skills, they are faced with confusion about what constitutes data literacy (Bhargava et al., 2015). Firms recognized as leaders in building data-driven cultures recognize the need and invest heavily in comprehensive, ongoing training (Brown et al., 2019; Landi, 2019), while many firms recognize the need, but struggle with knowing how to start (Columbus, 2014).
The findings of significantly lower scores for F3 across student and workforce participants, suggest that basic data wrangling skills are generally lacking. F1 scores suggest participants feel confident making data-driven decisions from provided data. F2 suggests that participants feel confident collecting data from various sources, but generally lack the skills to clean, model, and enrich data, limiting their ability to conduct meaningful analysis. This finding suggests academics and practitioners should create learning experiences to enhance basic data wrangling skills to enhance overall data literacy skills.
Data analytics involves a continuum of skills ranging from basic knowledge of data to advanced analytics involving complex algorithms, which numeracy and statistical skills (Kirby, 2021). The terms “data analysis” and “data analytics” are often used interchangeably, yet they refer to opposite ends of the skills continuum, which leads to confusion and contributes to the lack of traction for organizations seeking to improve skills. Individuals in possession of the full range of skills are highly desirable yet in short supply. Organizations must define the relevant data literacy skills for individuals to be successful in their roles, then create learning pathways to develop these individuals. The BDLI provides a starting point for educators and company trainers.
In developing the briefer data literacy instrument, we recognize the limitation that this research does not seek to measure the results of data literacy on a particular outcome. Future research should examine how data literacy impacts results across different industries and disciplines.
A composite study uses the same instrument data as used in Study 1 and Study 2 but removes the three weakest questions. We recognize that presenting the 12-question instrument to an audience may be preferred, it seems reasonable that an individual’s response to each question is independent and not influenced by the remaining questions.
Footnotes
Appendix A
Acknowledgements
The authors would like to acknowledge Dr. Ravi Nath, who was instrumental in the early work on this effort. Dr. Nath passed in January 2024.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Compliance with ethical standards
All procedures performed in studies involving human participants were in accordance with the ethical standards of the institutional research committee and with the 1964 Helsinki Declaration and its later amendments or comparable ethical standards.
Informed consent
Informed consent was obtained from all individual adult participants included in the study.
