Abstract
The data revolution has increased data demands for leveraging on Big Data in the production of statistics. The paper assesses the adoption of Big Data in research institutes in Kenya. Data were collected from 64 data practitioners based in the 24 research institutes that have a mandate to produce and analyse official statistics. The paper establishes the risks and challenges of using Big Data in statistics, identifies the determinants of adoption of Big Data in statistics and validates the relevance of a Technology Adoption Model (TAM) for predicting the adoption. It is the conclusion that there are immense opportunities for Big Data in statistics if the associated risks and challenges are addressed and the identified key determinants prioritized to promote the adoption.
Keywords
Introduction
The advent of Big Data is expected to have a big impact on institutes for which the production and analysis of data and information is core business [1]. The National Statistical System (NSS) is composed of all institutes and units within a country that jointly collect, process and disseminate official statistics on behalf of a national government [2]. Traditionally, the official statistics produced by the research and statistical institutes have been based on national surveys and censuses. However, generation of statistics can benefit by leveraging on administrative and secondary data sources in the production of statistics [3]. These traditional data sources are costly, the cycle of data collection is often infrequent hence quickly becomes outdated, and the data are usually reported at the national level so that spatial variations across a country are not often captured [4].
The Kenyan Statistics Act, 2006 established the Kenya National Bureau of Statistics (KNBS) (
Advances in technology have led to a data revolution characterized by an explosion of the volume of data with corresponding widespread and growing demand for data [7]. To support this data revolution and improve the quantity, frequency, disaggregation, and availability of relevant statistics, there is a need to make use of new sources of data made possible by innovations in technology [8]. Big Data is a transformative tool with great potential to fill gaps in statistics and potentially be leveraged to reduce costs and improve the availability of data to monitor development goals [2, 9].
The opportunities, challenges, and risks of Big Data are set to disrupt the production of official statistics by research and statistical institutes [10]. The statistical system will indisputably feel the pressure from the alternative sources of data and hence need to re-evaluate the usability of those sources in providing the information that decision-makers need. However, in the absence of a partnership between Big Data and Statistics, official statistics risk obsolescence, since Big Data will become increasingly attractive to data users [11]. Moreover, without coordination, Big Data may add to the cacophony of data discrepancies. This paper sought to assess the adoption of Big Data in research institutes in Kenya, establish the risks and challenges of using Big Data in statistics, identify the determinants of adoption of Big Data in statistics and validate the relevance of the TAM-based model for predicting the adoption.
Literature review
Data revolution
Data revolution revolves around Big Data, open data, data infrastructures and their consequence [12]. The revolution is founded on the proliferation of Information Communication Technologies(ICTs) such as mobile, distributed, cloud computing, social media, IoT, which are all reconfiguring the production, circulation and interpretation of data hence the term Big Data. Data revolution can greatly improve the way we measure, track, and describe economic activity and hence influence government policies. Einav and Levin [13] have highlighted how the data revolution may prove useful in economics while [14] on the other have discussed the role of national statistical systems in the data revolution. The data revolution need to be exploited for sustainable development.
As of 2020, the number of active mobile subscriptions in Kenya stood at 59.8 million which translates to a 125.8% mobile penetration rate [15]. The 2019 Kenya Population and Housing Census [16] results showed that 20,694,315 individuals aged 3 years and above owned a mobile phone. 22.6% of individuals aged 3 years and above-used internet while 10.4% used a computer. The proportion of the population aged 15 years and above who searched and bought goods and services online was 4.3%. This usage of ICT is driving the data revolution in Kenya which is being characterized by an explosion of the volume of data with corresponding widespread and growing demand for data [7].
The United High-level Panel on the global development framework for the Millennium Development Goals (MDGs) states in their [7] report that better data and statistics will help governments track progress, make evidence-based decisions and strengthen accountability. In addition to governments, international agencies, civil society organizations, and the private sector should be in the data revolution. The revolution will draw on existing and new sources of data to fully integrate statistics into decision making, promote open access to, and use of, data and ensure increased support for statistical systems [7]. Countries and citizens should be able to monitor development progress, hold leaders to account, and promote sustainable development.
Characteristics and sources of Big Data
Big Data is the collective noun for all of the new digital data arising from digital activities. In a world where increasingly day-to-day activities depend on technology, it seems that just about everything we think or do is now potentially a source of data. Big Data is being generated from a wide range of activities and transactions that are carried out including spending and travel patterns, online search queries, reading habits, television, and movies choices, social media posts [17]. [18] expounds on the sources of Big Data that includes administration of government or private sector programs (administrative sources) such as bank, hospital visits, electronic medical and insurance records; commercial or transactional sources arising from the transaction between two entities such as credit card, mobile and online transactions; sensor networks sources such as climate sensors, satellite imaging, and road sensors; tracking device sources such as tracking data from the Global Positioning System and mobile telephones; behavioral data sources such as online page views; and opinion data sources, such as social media activity.
Big data have been popularly characterized by five V’s in the ICT literature, namely, Volume, Velocity, Variety, Veracity, and Vulnerability [19]. Volume – refers to the amount of information available in Big Data, has been advocated as a plus for official statistics. Velocity – refers to the speed and frequency at which the data is generated. Variety – refers to the different types of data available which are broadly described as structured and unstructured data. Veracity – refers to the quality, accuracy of the Big Data, and the value of the insight that may be generated from the data set. Vulnerability – refers to the risk of exposure of the Big Data set to cybersecurity threats or disclosure of individual’s information. Big Data can further be described using 3C’s namely crumbs, capacities, and community [20].
Big Data in official statistics
The digital trails from the day-to-day dependence on technology offers official statisticians rich opportunities to enhance or displace existing data sources or generate completely new statistics [18, 21]. [22] have reported on what official statistics can expect from Big Data by giving an overview of Big Data pilots and projects conducted in official statistics at national and international levels. They indicate that official statistics cannot expect Big Data to substitute actual data sources but rather supplement the production of certain statistics. Increasing computing power opens up a myriad of new statistical possibilities where digital data can be shared, cross-references and repurposed as never before. [18] have described the Big Data Flagship Project of the Australian Bureau of Statistics while [21] have reported on the experience from analyzing Dutch traffic loop detection records and Dutch social media messages. Other illustrations on the role of Big Data in official statistics are found in [8, 20, 23, 24].
Across the world, national statistical offices have big data projects ongoing. They are using exhaust data, digital content, and sensing data from web scraping, Google maps, call detail records, satellite, and Twitter. Some of the projects noted include infection prevention and control, developing water accounts, tourism monitoring, subjective wellbeing, movements across borders, and complementing the national agriculture census [25].
This paper aims to contribute to the ongoing and future discussion of the role of Big Data in official statistics and development.
Opportunities, challenges and risks of Big Data for official statistics
Using Big Data in statistics leads to increased coverage, level of granularity, frequency, timeliness and reduced costs [26]. [17] expounds that assuming access problems can be overcome then Big Data offers the potential to contribute to the measurement of statistics in several ways. [23] group the opportunities for Big Data in statistics into five categories: entirely replace existing statistical sources such as surveys, partially replace existing statistical sources such as surveys, provide complementary statistical information in the same statistical domain but from other perspectives, improve estimates from statistical sources, and provide completely new statistical information in a particular statistical domain.
[18] adds the following opportunities: sample frame or register creation which is identifying survey population units and/or providing auxiliary information such as stratification variables; linking to other data which is creating richer datasets and/or longitudinal perspectives; imputation of missing data items which is substituting for same or similar units; data confrontation which is ensuring the validity and consistency of survey data; and editing which is assisting the detection and treatment of anomalies in survey data. Big Data can enhance, strengthen and complement statistics by providing variables to stratify sample surveys better, improving sample survey estimates, compensate for nonresponse, checking estimates, improve frequency and timeliness of data releases and improve and provide more small-area estimates [27].
Technology Acceptance Model (TAM) Source: [33].
Legal and regulatory issues guiding statistics production, data access and partnership negotiations, related to terms of access, privacy protection, and supply continuity, technological and methodological feasibility, establishing suitability for purpose and dataset quality, and the internal technical, financial and technological resources of the organizations are the major challenges related to big data use for statistics. [10, 26] envision dimension, quality, time dependence, and accessibility as the four major challenges for Big Data usage in Official Statistics. Dimension has an impact on storage of data – which requires advanced data architectures that can be supported by recent technologies such as Hadoop, “NoSQL” databases e.g. HBase, Hive, Google BigTable, Cassandra, and MongoDB; and processing of data which is a hard task both in terms of feasibility and efficiency.
[28] categorize the challenges of Big Data into three main groups: data (challenges relates to the characteristics of the data itself), process (challenges encountered while processing the Big Data), and management (legal and ethical issues related to accessing data). Gaining access to the required Big Data for assessment, experimenting, trialing and adoption is a challenge [18]. [17] categorize the statistical and governance challenges presented by Big Data into legal, ethical, technical, and reputational.
Although some Big Data is produced by public agencies, much of the Big Data is presently generated by private companies such as mobile phone, social media, utility, financial and retail companies and are valuable commodities to these companies, either providing a resource that generates competitive advantage or constituting a key product. This data is generally not publicly available for official or public analysis in raw or derived forms [10]. Producers of Big Data gain their competitive advantage by maintaining their data locked and hence the accessibility challenge [29]. [30] have provided a critical analysis of Big Data challenges with a view to assisting organizations make robust investment decisions. Big Data requires development of new theoretical frameworks and methodologies for it cannot be adapted to existing ones. Representativeness of Big Data available for a population of interest is a major challenge [22].
The key risks relate to mission drift, reputation and trust, privacy and data security, access and continuity, fragmentation across jurisdictions, resource constraints and cut-backs, and privatization and competition [10, 31]. Privacy concerns severely hamper Big Data producers to share their private data which is sensitive and sometimes trade secrets [32]. Risks can be caused by poor data quality, inappropriate use of data and analytics and unexpected change in the Big Data environment. Inappropriate use of Big Data and Analytics risk underscores that domain knowledge and expertise is necessary to make decisions and not rely on software generated analytical facts which may neglect biases hence lead to false conclusions [24].
The study adopted the Technology Acceptance Model (TAM) (Fig. 1) that was developed by [33] as an extension of Ajzen and Fishbein’s (1980) Theory of Reasoned Action (TRA) which has been used extensively in many studies to evaluate the acceptance and use of technology. TAM evaluates acceptance of technology using perceived usefulness and perceived ease of use. Perceived Usefulness is the degree to which a person believes that using a particular system would increase your work performance. Perceived ease of use is the degree to which a person believes that using a particular system would be effortless. The two factors are influenced by external factors that may be either social (language, skills, and facilitating conditions), cultural, or political [34].
The TAM-based model (Fig. 2) was extended using the following constructs: Interpersonal/Peer influence, External influence, Compatibility, Self-Efficacy, Facilitating Conditions, Perceived Behavioral Control, and Subjective Norms.
Research model.
From the model adopted, the following are the determinants of the adoption of technology.
Social Influences – Derived from [35] Unified Theory of Acceptance and Use of Technology, Social influence means that in some cases people might use technology to comply with the mandates of others rather than their feelings and beliefs. Social influences can either be interpersonal or external.
H1:
Interpersonal influence will have a significant positive effect on Subjective norms.
H2:
External influence will have a significant positive effect on Subjective norms.
Perceived Usefulness – Defined as the degree to which a user believes that using a specific system would enhance job performance [33].
H3:
Perceived usefulness will have a significant positive effect on Attitude towards the use of Big Data in statistics.
Perceived Ease of Use – Defined as the degree to which a user believes that using a particular system would be effort-free [33].
H4:
Perceived ease of use will have a significant positive effect on Attitude towards the use of Big Data in statistics.
Compatibility – This is an important dimension of the innovation diffusion theory. In [36] meta-analysis of innovation, they found that innovation was more likely to be adopted when it was compatible with the individual’s job responsibilities and value system.
H5:
Compatibility will have a significant positive effect on Attitude towards the use of Big Data in statistics.
Self-Efficacy – Self-efficacy is defined as the belief that one can perform a particular behavior.
H6:
Self-efficacy will have a significant positive effect on Perceived Behavior Control.
Facilitating Conditions – Derived from [35] Unified Theory of Acceptance and Use of Technology, these conditions that may affect the adoption of technology include resource factors, technology factors, availability of training and provision of support, and external controls such as policies, regulations, and legal environment.
H7:
Facilitating conditions will have a significant positive effect on Perceived Behaviour Control.
Perceived Behavioural Control – According to [37], perceived behavioral control reflects belief regarding access to the resources and opportunities needed to affect behavior. It encompasses facilitating conditions and self-efficacy.
H8:
Perceived behavioral control will have a significant positive effect on the Intention to use Big Data in statistics.
Attitude Towards Use – Defined as the mediating effective response between usefulness and ease of use beliefs and intentions to use technology [33].
H9:
Attitude will have a significant positive effect on the Intention to use Big Data in statistics.
Subjective Norms – This is a person’s belief that most of the others who are important to him think he should (or should not) perform the behavior in question [38]. Social influence is a determinant of the behavior of subjective norms since in accepting a technology, users tend to be affected by external and interpersonal/peer influences.
H10:
Subjective norms will have a significant positive effect on the Intention to use Big Data in statistics.
Intention to Use Technology – Defined as the user’s likelihood to use a tool or a technology [33]. Behavior is determined by behavioral intention as long as an individual agrees to perform a behavior [37]. A prospective user’s overall attitude toward using a given technology is an antecedent to intentions to adopt.
Research design
A research strategy is designed specifically for the organisation to manage bias and ensure reliability, validity, and normality [39]. The research model depicted in Fig. 2 was used to frame the study and guide on data collection. The research study used a deductive approach achieved by the use of a descriptive survey where quantitative data was collected and statistically analyzed to draw conclusions based on the research model. The quantitative design drew from a positivist perspective [40].
Data collection
This study targeted practitioners dealing with data and the generation of statistics in research institutes in Kenya. The study targeted four data practitioners from each of the 24 research institutes. The data was collected using the survey technique [41, 42] whereby self-administered questionnaires were sent to sampled respondents via email which was most convenient given the existing COVID-19 pandemic. Where the questionnaire failed to work, telephone interviews were conducted. The questionnaire was categorized into eleven sections where the first section contained demographic attributes of the respondent whereas the other ten sections sought to answer the dimensions of the TAM-based model used: Interpersonal/Peer Influence (IP), External Influence (EI), Perceived Usefulness (PU), Perceived Ease Of Use (PEOU), Compatibility (CP), Self-Efficacy (SE), Facilitating Conditions (FC), Perceived Behavioral Control (PBC), Attitude Towards Use (ATU), Subjective Norms (SN), and Intention To Use (ITU). Specific questions for each dimension were asked and answers were selected from a five-point Likert scale with a range: Strongly Agree
Validity and reliability
Reliability and validity are the key indicators of the quality of a research instrument and hence the process of developing and validating an instrument is majorly focused on reducing the error of the measurement process. [43] elaborate that reliability estimates evaluate the internal consistency of measurement instruments, interrater reliability of instrument scores, and stability of measures whereas validity is the extent to which the interpretations of the results of a test are warranted. Reliability was assessed using composite reliability (CR) while factor loadings and average variance extracted (AVE) were used to assess convergent validity [44].
Data analysis
Data cleaning was done to remove any errors and inconsistencies then coded before being imported to Stata software for analysis. Descriptive statistics analysis was carried out which included measures of central tendencies, measures of dispersion, and measures of frequency. Reliability was assessed using composite reliability (CR) while factor loadings and average variance extracted (AVE) were used to assess convergent validity [44]. The path analysis method of structured equation modeling technique was used for validation of the conceptual model used. Charts, graphs, and tables were used to represent the data.
Ethical consideration
The rights of the participants as highlighted by [42] i.e. not to participate, to withdraw, to give informed consent, anonymity, and confidentiality were adhered to. The participants involved were fully briefed about the purpose of the research in advance and responding to the study was optional. The anonymity of the participants was achieved by not collecting their names and locations. The research institution’s identity was anonymized. The data collected was handled safely and securely to enhance confidentiality. The data collected was used for academic purposes only. Additionally, the participants were protected from incurring any cost when undertaking the survey.
The responsibilities of the researcher as highlighted by [42] i.e. no unnecessary intrusion, behave with integrity, follow appropriate professional codes of conduct, no plagiarism, and be an ethical reviewer were adhered to. The privacy of the respondents was respected and all literature quoted was fully credited.
Big Data usage in research institutes.
Major Big Data Risks.
A total of 96 respondents were targeted for this study out of which 64 responses were received, a response rate was 67%. A response rate of 60% and above is considered good [45].
Respondents demographics
The demographics of the respondents are shown in Table 3.
Big Data usage in research institutes, risks and challenges
Forty-eight per cent of the respondents confirmed that their institutes have a Big Data Strategy in place. Four questions to measure whether Big Data has been adopted at the organization level were asked and the results are summarized in Fig. 3. 91% of the respondent’s institutes foresee a role of Big Data in the execution of their work indicating prioritization and incorporation of Big Data in their operations.
Demographic characteristics
Demographic characteristics
92% of the respondents have existing potential Big Data sources that they can use in their work. This can be deduced to mean that there exists a lot of unutilized data that can be used in the analysis and generation of statistics. 94% of the respondents think that Big Data has the potential to alleviate data challenges in their institutes. This positive response shows that the generation of statistics using Big Data will be adopted easily since the data practitioners recognize its potential usefulness. 94% of the respondents agree with the notion that Big Data can be used to supplement official statistics thus indicating that the practitioners in the data sector welcome the rise of new data sources to complement the traditional sources.
Major Big Data challenges.
Figure 4 details the sentiments of the respondents regarding the major challenges highlighted. In concurrence to [10] observations, the most prominent risks were: inconsistent access and continuity reported (69%); privacy breaches and data security (66%); resource constraints and cut-backs (55%); and resistance of Big Data providers and populace (47%).
Big Data challenges
Figure 5 details the respondent selection to the question regarding the Big Data challenges. These challenges had been adopted from [10] and submitted to the respondents for validation; 70% of the respondents reported legal and regulatory issues as major challenges. 46% percent reported gaining access to data which was proportional to gaining access to associated methodology and metadata. This signifies that data sharing is a major challenge. Data quality challenges (55%) arise due to its high veracity, uncertainty, error, bias, reliability, and lack of calibration. Other challenges are establishing suitability of purpose (39%), institutional change management (31%) and ensuring inter-organizational collaboration (22%).
Conclusion
With Kenya and the rest of developing countries lagging behind with the number of statistical indicators being tracked, using Big Data is a necessity for evidence-based decision making since it is generated at a lower cost, greater frequency and with a wider spatial distribution. Research institutions which are hubs for innovations and inventions must embrace Big Data complement their functions.
This paper established that data professionals in research institutes agree that Big Data can complement traditional sources of data to generate statistics and are ready to adopt it. This calls for the training and development of employee to acquire the requisite skills.
The challenges of legal and regulatory issues, gaining access to data and associated methodology and metadata, and establishing Big Data quality were noted to be the main challenges hindering the adoption and should be resolved. This calls for policy makers to create an enabling legal and regulatory environment. Data sharing polices and agreements should be developed to handle the risk of inconsistent access of Big Data. Enforcement of non-disclosure agreements and clear definition of personal data should be done to deal with privacy breaches.
The study established that external influence and subjective norms, perceived usefulness, compatibility, attitude towards use, and self-efficacy are the key factors influencing adoption of Big Data in Statistics.
This paper recommended that research institutes undertake a radical shift in statistical methodology for Big Data to gain ground in statistics. The current statistical methodologies of sampling and analysis should be improved to handle Big Data. A quality framework and quality criteria need to be designed or adapted for the various types of Big Data and statistical products that are based on these data [22]. Sensitization and capacity building of all stakeholders on the opportunities of Big Data for statistics should be done to improve the attitude towards use which is a key determinant of the use of Big Data in statistics. Facilitating conditions such as legal and regulatory issues should be resolved by full implementation of the data protection act of 2019 and allocation of enough resources to Big Data projects. Training on new methodology and tools to handle Big Data should be enhanced to enhance self-efficacy. Advocacy should be enhanced since external influence and subjective norms play a significant role in the adoption.
Footnotes
Acknowledgments
The authors acknowledge the contribution of Daniel Orwa of the Department of Computing and Informatics, University of Nairobi and Mary Karumba of the State Department of Planning, Kenya for their critical review of this research work, and the staff of Kenya National Bureau of Statistics who, in one way or the other, provided the required research data.
Appendix
Summary of model descriptive statistics
Construct
Mean
SD
Skewness
Kurtosis
Interpersonal/Peer Influence
3.8177
0.4279
0.431
2.4016
(IP)
External Influence (EI)
3.8802
0.4337
2.5353
Perceived Usefulness (PU)
4.1367
0.4603
2.7686
Perceived Ease Of Use
3.9014
0.4393
3.1864
(PEOU)
Compatibility (CP)
4.0305
0.4533
4.0649
Self-Efficacy (SE)
3.5797
0.4052
2.0911
Facilitating Conditions (FC)
3.3268
0.3811
2.0071
Perceived Behavioral Control
3.3704
0.3800
2.2295
(PBC)
Attitude Towards Use (ATU)
4.3261
0.4843
2.7722
Subjective Norms (SN)
3.7773
0.4237
2.1069
Intention To Use (ITU)
4.2210
0.4721
3.2204
Reliability and factor loadings
Constructs/measurement item
Standardized loading
CR
AV
Interpersonal/Peer Influence (IP)
Interpersonal1
0.8989
Interpersonal2
0.4964
Interpersonal3
0.7416
External Influence (EI)
Externalinfluence1
0.5025
Externalinfluence2
0.6659
Externalinfluence3
0.6709
Perceived Usefulness (PU)
Pu1
0.6578
Pu2
0.8127
Pu3
0.7525
Pu4
0.6661
Perceived Ease Of Use (PEOU)
Peou1
0.5547
Peou2
0.5663
Peou3
0.7900
Peou4
0.7258
Compatibility (CP)
Compatibility1
0.7915
Compatibility2
0.8450
Compatibility3
0.8327
Self-Efficacy (SE)
Selfefficacy1
0.7811
Selfefficacy2
0.9263
Selfefficacy3
0.8398
Facilitating Conditions (FC)
Facilitating1
0.8075
Facilitating2
0.7929
Facilitating3
Perceived Behavioral Control (PBC)
Pbc1
0.6120
Pbc2
0.9428
Pbc3
0.6395
Attitude Towards Use (ATU)
Attitude1
0.8053
Attitude2
Attitude3
0.7384
Attitude4
0.8047
Subjective Norms (SN)
Sn1
0.7435
Sn2
0.8419
Sn3
0.7036
Intention To Use (ITU)
Intention1
0.6963
Intention2
0.8455
Intention3
0.6032
Correlation matrix and roots of AVEs
IP
EI
PU
PEOU
CP
SE
FC
PB
AT
SN
IT
IP
EI
0.6132
PU
0.5412
0.5665
PEOU
0.2312
0.3797
0.4666
CP
0.5148
0.6052
0.6950
0.5352
SE
0.2739
0.3399
0.3865
0.5715
0.4579
FC
0.4157
0.4405
0.1516
0.2457
0.1642
0.4589
PB
0.3662
0.4265
0.3553
0.6107
0.4152
0.8440
0.6263
AT
0.5085
0.5582
0.6098
0.2201
0.8232
0.2514
0.0774
0.2233
SN
0.6871
0.5691
0.5315
0.3427
0.5611
0.5135
0.5647
0.6069
0.5915
IT
0.6739
0.5213
0.6992
0.2685
0.7786
0.4230
0.3216
0.4467
0.7396
0.7271
