Abstract
Science, Technology, Engineering, and Mathematics student success is an important topic in higher education research. Recently, the use of data analytics in higher education administration has gain popularity. However, very few studies have examined how data analytics may influence Science, Technology, Engineering, and Mathematics student success. This study took the first step to investigate the influence of using predictive analytics on academic advising in engineering majors. Specifically, we examined the effects of predictive analytics-informed academic advising among undeclared first-year engineering student with regard to changing a major and selecting a program of study. We utilized the propensity score matching technique to compare students who received predictive analytics-informed advising with those who did not. Results indicated that students who received predictive analytics-informed advising were more likely to change a major than their counterparts. No significant effects was detected regarding selecting a program of study. Implications of the findings for policy, practice, and future research were discussed.
Introduction
Retention and completion in the Science, Technology, Engineering, and Mathematics (STEM) disciplines have been highlighted as an important issue in higher education in the United States. According to a recent report, about 48% of bachelor’s degree seeking students in STEM fields had left their original field of study within 6 years (Chen, 2013). Among those who left, approximately one half changed majors to a non-STEM field while the remaining two thirds dropped out of college without earning a degree or certificate (Chen, 2013). Graduating a sufficient number of students in STEM majors has been a priority nationwide (National Science and Technology Council, 2013; National Science Foundation, National Center for Science and Engineering Statistics, 2017) . To address this priority, higher education institutions need to strategize to retaining current STEM students in addition to recruiting new students into STEM majors. One particular student group of interests for this discussion is the first-year college students in STEM who have not declared their program of study.
Undecided students are often studied in higher education research focused on student retention and persistence. There is considerable diversity in this population (Gordon & Steele, 2015). For example, undecided students may refer to those who have not declared a major upon matriculation or those who have chosen a major but not selected a particular program of study. In this study, we focused on undeclared first-year, first-term engineering students. These students made the initial decision of enrolling in an engineering department but did not specify a program of study. They can move forward to selecting a program of study within engineering or changing to a nonengineering, or non-STEM major. It is imperative to identify and utilize effective strategies in higher education to promote undeclared engineering students’ academic success. One specific strategy is incorporating data analytics into academic advising which would influence students’ major and program of study selection.
Data analytics refers to collecting, sharing, and using data to identify patterns and trends and to provide evidence for decision-making (Barneveld, Arnold, & Campbell, 2012). It represents an emerging and critical topic in higher education administration and student affairs (Tinto, 2014). Although the use of data analytics in higher education is still in its infancy, various applications have started to provide observable outcomes (Dahlstrom, 2016). One example is using data analytics to develop statistical models that predict academic outcomes. This type of data analytics is called predictive analytics. In particular, using data generated from predictive analytics would allow advisors and faculty members to forecast the academic outcome of a student based on their current academic performance and progress. Such evidence-based forecasting can complement advisors’ experience and instinct to provide more informed academic advising, which would lead to more desirable academic outcome. As one of the earliest attempts, we are interested in empirically examining the influences of data analytics (predictive analytics) on STEM student success.
Literature Review
Choice of Major Among College Students
The role of choice of major in student success
Various studies have highlighted the importance of choice of major in college students’ success. In general, matriculating with declared majors is favorable for students’ long-term persistence and academic success (Kreysa, 2006; Leppel, 2001). Besides studying the influences of declared major upon matriculation, previous studies also examined the effect of changing a major versus not changing a major. Researchers indicated that undergraduates who changed a major really knew what they want in terms of future career; thus, they are more likely to graduate and continue to pursue their career goals (Foraker, 2012; Kreysa, 2006; Micceri, 2001; Murphy, 2000).
In terms of the timing of changing a major, some researchers have argued that it is not detrimental to delay on the decision (Cueso, 2005). However, others have posited that delaying decisions can have a negative impact on academic success and financial liability (Foraker, 2012; Jenkins & Cho, 2012; Sklar, 2014). For instance, Foraker (2012) concluded that changing a major (or declare a major as undeclared students) early in the college journey indicated that the student made a well-informed and refined decision. However, if the decision of changing a major or declaring a major was made later (i.e., after the second year), students might have lower grade point average (GPA), lower graduation rates, and need longer time to completion (Foraker, 2012). Under the similar rationale, students who have chosen a major but did not select a program of study should enter a program of study as soon as possible for better academic outcomes (Jenkins & Cho, 2012)
Influential factors of choice of major
Studies in and outside the field of higher education have generated fruitful findings regarding factors that contribute to students decision of choice of major. For example, in the field of vocational psychology, career indecision has been studied since 1950s (Gordon & Steele, 2015; Miller & Rottinghaus, 2014). Career indecisiveness is closely related to being undecided in college (Gordon & Steele, 2015; Kreysa, 2006; Sandler, 2000). Potential reason for vocational/career indecision include vocational identity, career self-efficacy, and career-related anxiety (Brown & Rector, 2008).
Studies in higher education also highlighted the critical role of psychological factors (Allen & Robbins, 2008; Arcidiacono, 2004; Eccles, 1987; Stinebrickner & Strinebrickner, 2011). Specifically, Arcidiacono (2004) determined that choice of major was influenced mostly by how students perceived their ability to complete coursework in the major they had chosen. Other higher education studies also confirmed the critical role that academic self-efficacy played in regard to choice of major (Allen & Robbins, 2008; Eccles, 1987; Stinebrickner & Strinebrickner, 2011).
Further, academic and social integration are important factors that were included in well-known theoretical models for college student retention and persistence (Cabrera, Castaneda, Nora, & Hengstler, 1992; Cabrera, Nora, & Castaneda, 1993; Pascarella & Terenzini, 1980; Tinto, 1975). In these models, academic and social integration were often measured by students’ interactions with faculty members, peers, classmates, and friends (Cabrera et al., 1993; Pascarella & Terenzini, 1980). When predicting students’ behaviors regarding major choice, these measures were repeatedly highlighted as significant predictors (Cohen & Hanno, 1993; Mauldin, Crain, & Mounce, 2000).
In addition, demographic and background characteristics also influence choice of major. Previous studies have demonstrated the significant influences from demographic characteristics such as gender (Turner & Bowen, 1999; Zafar, 2013), ethnicity and race (Arcidiacono, Aucejo, & Spenner, 2012; Maple & Stage, 1991; Porter & Umbach, 2006), parental influence (Astin, 1993), and expected income (Berger, 1988; Boudarbat, 2008).
While the psychological factors, college experiences (i.e., academic and social integration), and demographic characteristics are part of the overall picture, the actual academic performance is another critical piece to add (Cabrera et al., 1993; Foraker, 2012). As more data about an individual’s academic performance become available, academic advisors can better help students to predict their future outcomes within a major and lead them to major choice/change decision.
Retention in STEM and Engineering
STEM student retention has been an important and critical issue for higher education researchers and practitioners for the last decades (i.e., Starobin & Laanan, 2008; Watkins & Mazur, 2013; Xu, 2018). Many significant predictors of retention and persistence across all majors were proved critical for STEM student as well. For example, self-efficacy has a critical role in influcing academic progress of STEM students (Pajares, 1996). For engineering students, academic self-efficacy and self-concept as an engineer are especially critical (Eliot & Turns, 2011; Lent, Brown, & Larkin, 1984).
Further, academic integration also contribute to STEM student retention and persistence. For example, Xu (2018) found that STEM students were less likely to leave the major if they felt positively about teaching quality, faculty support, and academic advising. Ferrare and Lee (2014) found that involving in study groups and taking more STEM credits during the first year significantly predicted STEM persistence.
Lastly but not the least, students with different demographic and background characteristics show different patterns in STEM retention. For example, women and racial minorities have been underrepresented in STEM majors (National Science Foundation, National Center for Science and Engineering Statistics, 2017). Factors that are critical in retention for these students are different from those are important for White males (Maple & Stage, 1991).
Data Analytics in Higher Education
Data analytics refers to using data to identify patterns and trends and to provide insights that guide decision-making appropriately (Barneveld et al., 2012). It has been a valuable asset in business- and health-related fields (Hersh, 2002; Ngai, Xiu, & Chau, 2009). In higher education, however, data analytics is still in its infancy (Tinto, 2014). Although many higher education institutions are collecting data, these data were mainly used for credentialing and reporting purposes rather than generating strategies and solutions to achieve institutional goals (Bichsel, 2012). It is imperative to apply data analytics in the area of promoting students’ success.
Among various data analytics approaches, predictive analytics is an emerging area pertaining to student success. Predictive analytics highlights using statistical analysis tools and techniques to undercover relationships and patterns (Barneveld et al., 2012). The ultimate goal of predictive analytics is to predicting student behaviors and potential outcomes in the future (Barneveld et al., 2012). One way to implement predictive analytics might be construct and utilize a predictive mechanism (Campbell, Deblois, & Oblinger, 2007). Such mechanism can be a system, a software, or a statistical model that generate prediction based on student data collected from registrar’s office, financial aid offices, residential halls, and more. Although there are several institutional reports on implementing predictive analytics tools (i.e., Barber & Sharkey, 2012; Denley, 2014), there is a lack of knowledge with regard to its impact on student outcomes.
Purpose and Research Questions
This study aims at exploring whether utilizing predictive analytics in academic advising can significantly impact student success in an engineering department at a research intensive, public 4-year university located in the Midwest. In this university, predictive analytics was used to facilitate advisors’ interaction with undeclared first-year, first-term engineering students. Under the guidance of advisors, these students either were encouraged to select a program of study within engineering as early as possible or explore alternate majors. The latter option is for those who were experiencing academic challenges. In this study, we are interested in determining the impact of predictive analytics during academic advising on student behaviors. We investigated whether the use of predictive analytics in advising had any influences on two student behaviors: (a) continuing to seek a program of study in engineering or selecting a program of study and (b) leaving engineering or changing a major. The following research questions guided this study:
Are there any observable demographic and academic differences between students who received predictive analytics-informed academic advising and those who did not? Which factors and variables have significantly predicted undeclared engineering students’ behaviors of changing of major after the first term? Which factors and variables significantly predicted undeclared engineering students’ behaviors of selecting a program of study after the first term? Did predictive analytics-informed academic advising have an impact on undeclared engineering students’ behaviors in terms of changing a major after the first term? Did predictive analytics-informed academic advising have an impact on undeclared engineering students’ behavior in terms of selecting a program of study after the first term?
Theoretical Framework
The theoretical framework in this study is threefold. Specifically, we adopted Astin’s Input-Environment-Output (I-E-O) model as the main theoretical framework. We also considered Tinto’s academic integration theory and the role of self-efficacy to complement the specific characteristics of undeclared engineering students.
The I-E-O Model
The primary theoretical framework of this study is Astin’s (1993) I-E-O model. In the I-E-O model, students bring their inherited demographic characteristics and previous academic experiences in high school as input to the postsecondary institutions. As students are developing psychologically, socially, and academically, adjustments and adaptions are consequently made to their higher education environment. If we perceive the choice of major as one of the transitional outputs, then we must highlight both the inputs and influences of socioenvironmental factors in higher education institutions (Astin, 1993; Tinto & Pusser, 2006).
Academic Integration Theory
Academic integration theory (Tinto, 1975, 1987) also provides an explanation of persistence and choice of major. Tinto (2005, 2014) emphasized the importance of goal commitment, academic integration, and intellectual development. The experiences with academic advisors are important parts of students’ academic integration. While a new academic advising tools may play a role with respect to an undeclared engineering students’ academic integration, we also intend to consider other academic integration aspects such as interaction with faculty members, peers, and other student affairs professionals.
Self-Efficacy Theory
The extent to which students are able to formulate, adapt, and make decisions in their academic environments is linked to self-efficacy (Bandura, 1993; Zimmerman, Bandura, & Martinez-Pons, 1992). Self-efficacy is manifested in cognitive, motivational, affective, and selection processes. Previous studies showed its critical role in engineering students (Eliot & Turns, 2011; Lent et al., 1984) and its positive influence on choice of major (Arcidiacono, 2004, Pajares, 1996). In this study, we view self-efficacy as an important predictor of either persisting in engineering major or changing a major.
In sum, we propose a model (Figure 1) that includes analytics-informed advising as a key environmental variable in addition to other input variables (i.e., demographic variables and precollege academic characteristics) and environmental variables (i.e., academic self-efficacy, integration, first-term experiences, and academic advising).

Model of influences on change of major and selection of program of study.
Methods
This study utilized a quasi-experimental design with observational data. We selected first-year, first-term, undeclared engineering students as the targeted population. The treatment was defined as the predictive analytics-informed academic advising.
Data Source
The data in this study were collected from a large, research-intensive, public 4-year university, Large Midwest Research University (or LMRU, pseudo-name). The data set was created by integrating two institutional data sources at LMRU: student information system (SIS) and Mapworks® data set. Specifically, the SIS database provided data regarding students’ demographic characteristics, preenrollment academic characteristics, academic engagement activity, and first-term academic performance. The Mapworks® data set contained data about students’ academic self-efficacy, academic integration, and social integration. The Mapworks®, or the Make Achievement Possible, is a comprehensive student success and retention system developed by Ball State University and Skyfactor (formerly EBI MAP-Works). Institutional practitioners can use the survey data from Mapworks® to identify at-risk students and employ early intervention to enhance students’ success and retention (Skyfactor, 2014). The validity and reliability of the Mapworks® survey has been well established and tested through its solid theoretical grounding and multiple years of implementation (Skyfactor, 2014). The Mapworks® has been adopted by LMRU since 2008, which is often disseminated in the third or fourth week of the fall term to all freshmen. The participation rates in the Mapworks® survey have exceeded 80% at LMRU since 2012.
Sample
We included students enrolled between the Fall 2012 and Spring 2016 academic semesters in the College of Engineering at the LMRU. We focused only undeclared first-year, first-term engineering students during the selected semesters. Using student ID numbers, the SIS data records were matched with Mapworks® responses. Student identifiers were removed prior to analysis. Because the Mapworks® survey instrument was revised in the Fall 2012, we did not utilize data prior to 2012 to avoid inconsistency.
Starting in the fall of 2015, the academic advising in the College of Engineering at LMRU started to use a predictive analytics software. Students enrolled before this implementation (Fall 2012–Spring 2015) were deemed as the control group (n = 702), representing 51% of all undeclared first-year, first-term engineering students between Fall 2012 and Spring 2015. Students who were exposed to the academic advising informed by the predictive analytics software belonged to the treatment groups (n = 125), representing 26% of all undeclared first-year, first-terms engineering students in Fall 2015 and Spring 2016. Thus, the total sample consisted of 827 students between Fall 2012 and Spring 2016.
Variables Used in This Study
Dependent variables
There are two dependent variables in this study: “change of major” and “selection of program of study.” These variables were collected from the SIS data set based on students’ status on the 10th day of class in spring semester. These two dependent variables were dummy coded. For “change of major,” a “1” represented that a student changed majors and a “0” represented the student did not. Similarly, in the “selection of program of study” variable, a “1” means a student has selected a program of study and a “0” means the student did not.
Independent variables
The following independent variables were selected from SIS and Mapworks®. The SIS variables included age, gender, ethnicity, ACT composite score, ACT math score, high school rank, first-term credits attempted, first-term credits completed, first-term GPA, learning community memberships, and honors program member. The Mapworks® variables included commitment to completing degree, math self-efficacy, academic self-efficacy, and institutional satisfaction. In addition, we also included three factors from Mapworks®: academic integration, social integration, and academic self-efficacy. Coding for these variables are included in Table 1.
Independent Variables Used in This Study.
Note. GPA = grade point average; LMRU = Large Midwest Research University.
Data Analysis
Descriptive and comparative analysis
To answer the Research Question 1, group differences were examined using descriptive analysis and t tests. We compared students who received predictive analytics-informed academic advising with students who did not. We examined a variety of variables referring to students’ demographics characteristics, preenrollment academic characteristics, academic engagement activity, and first-term academic completion outcomes. We also compared the two groups in terms of academic self-efficacy, academic integration, and social integration.
Logistic regression models
Next to answer the Research Questions 2 and 3, binary logistic regression models were constructed. Our intention was to explore which and to what extent variables from SIS and Mapworks® were significant predictors of the two dependent variables. Two logistic regression models were developed. The following regression equations describe the models
1
:
To assess the model fit, four statistical indicators were examined: −2log likelihood, pseudo-R2, Hosmer–Lemeshow statistic, and classification table. The logistic regressions helped us guide the selection of variables in constructing subsequent propensity score estimation other than addressing the research questions.
Propensity score matching
To address Research Questions 4 and 5, we utilized propensity score matching (PSM) to determine the impact of predictive analytics-informed academic advising on student behaviors. Propensity score analysis is an appropriate method for this study because students were chosen nonrandomly into a treatment group versus a control group. The PSM technique is an effective way to control selection biases, which allowed us to form a quasi-experimental design in this study (Rosenbaum & Rubin, 1983; Rubin, 1997). Four phases of the PSM analysis were conducted: (a) estimating the propensity score, (b) matching the treatment and control groups, (c) checking for balance, and (d) estimating the treatment effect. First, the propensity score was estimated based on a probit regression model that reflected the probability of a student receiving the treatment. We established the model based on the results of logistic regressions in previous analyses. Next, we utilized the nearest neighbor matching techniques to match treatment group students with control group students who have the closest propensity scores. 2 Other matching methods such as stratification, kernel, and radius were used to confirm the impact of the treatment. Following the matching, a series of t tests were conducted to ensure balance (i.e., no significant differences) between the groups. The last step was to estimate the effect of treatment (i.e., predictive analytics-informed advising). We estimated the effect of the treatment by computing group differences in the outcome variables using the average treatment on the treated (ATT) measurement. A bootstrap repetition was utilized to calculate standard error of the treatment effect 3 (Caliendo & Kopeinig, 2008; Lechner, 2002).
Limitations
There are two limitations that we note. First, the sample is predominantly male (78%) and White (83.4%). Readers should be cautious about applying findings of this study to different student groups. Second, this study is limited by the data collected. Specifically, additional variables outside of the SIS and Mapworks data may also meaningfully contribute to the models. For example, we are unable to capture the quality and nuances of interactions between academic advisors and students. This may leave room for future research.
Results
Descriptive and Comparative Analysis
Descriptive and t test results revealed observable differences between the treatment and control groups. The control group had more males (79.3%) than the treatment group (70.4%). Both groups were mainly consisted of White individuals (85.0% and 81.6% for treatment and control groups, respectively) and of 17 and 18-year-olds.
Further, more students in the treatment group were in a learning community (p < .001) and participated in honors programs (p < .05). In addition, students in the treatment group reported higher academic self-efficacy and academic integration (p < .05). Conversely, the students in the control group showed significantly higher first-term credits attempted (p < .001). More between-group comparisons for the variables are presented in Table 2.
Descriptive and Comparative Statistics of Control and Treatment Groups.
Note. SD = standard deviation; GPA = grade point average; LMRU = Large Midwest Research University.
*p<.05. ***p < .001. ns = not significant.
To examine the patterns in student behaviors with regard to academic major, a descriptive bar chart was constructed to reveal the trends over the years with respect to change of major, selection of program of study, withdrawal from the institution, or remaining in an undeclared status. From the chart, it can be seen that the selection of program of study was at a higher percentage for the treatment group than for the control group. In the year of the treatment (i.e., 2015–2016 academic year), there was 8% increase in selection of program of study. This shift is also reflected by fewer students remaining in the undeclared status (i.e., almost 10% less than the preceding year).
Logistic Regression
To address Research Questions 2 and 3, two logistic regression models were constructed. The first logistic model was to determine significant predictors of changing a major. The logistic model exhibited a good model fit, χ2(8) = 6.608, Nagelkerke’s pseudo-R2 = 0.164, area under the receiver operating characteristic = 0.769, in predicting the odds of changing major with an overall classification success rate of 72.6%.
The logistic regression model (Table 3) revealed the following significant predictors for changing a major: first-term credits attempted, first-term credits completed, and ACT composite score. Specifically, for every additional three-credit course attempted, students were 3.74 times more likely to change majors. For every additional three-credit course completed, students were 3.68 times less likely to change major. In addition, for each additional point earned on their ACT composite score, students were 1.10 times less likely to change major. Finally, for every unit higher in academic self-efficacy, students were 1.66 times less likely to change a major.
Significant Predictors for Change of Major in Logistic Regression.
Note. n = 827. SE = standard error.
*p < .05. ***p < .001.
Similarly, the second logistic regression model was constructed to determine significant predictors for the selection of program of study. This model also exhibited a good model fit, χ2(8) = 6.480, Nagelkerke’s pseudo-R2 = 0.058, area under the receiver operating characteristic = 0.611, with an overall classification success rate of 58.9%.
This model (Table 4) revealed the odds of the selection of program of study were most influenced by first-term credits completed and academic integration. Specifically, for every additional three-credit course completed in the first term, students were 3.366 times more likely to select a program of study (or, for every one additional credit completed in students was 1.122 times more likely to select a program of study). In addition, for each unit increase in academic integration, students were 1.314 times more likely to select a program of study.
Significant Predictors for Selection of Program in Logistic Regression.
Note. n = 827. SE = standard error.
***p < .001.
Propensity Score Results
PSM was used to understand the impact of predictive analytics-informed academic advising. Two probit regression models were constructed to compute the propensity scores for estimating the treatment effects on changing a major and selecting a program of study, respectively. The two logistic regression models informed the selection of variables in the probit regression models. For the PSM on change of major, we included the following independent variables in the probit model: age, gender, ACT composite score, high school rank, first-term credits attempted, first-term credits completed, learning community memberships, and honors program member, academic self-efficacy, academic integration, and social integration. The probit regression model (Table 5) resulted in a moderately strong goodness of fit (McFadden’s pseudo-R2 = 0.7085). Next, the nearest neighbor matching algorithm was utilized to match control and treatment students. In particular, each treatment group student was matched with one control group student who had the closest propensity score. 4 Subsequent t tests were conducted on each variable using the mean propensity scores to ensure the balancing property between the control and treatment groups was satisfied. No significant differences were observed between groups for the selected variables following the matching. After matching, the treatment (i. e., analytics-informed academic advising) demonstrated a significant impact on students’ behavior of changing a major. The ATT value demonstrated that students who received analytics-informed advising were significantly more likely to change majors by 6.8%. We confirmed this significant effect on changing a major by using three other matching methods (i.e., stratification, kernel, and radius), as shown in Table 6.
Probit Regression for Change of Major.
Note. n = 827. SE = standard error.
***p < .001.
ATT Estimation With Bootstrapped Standard Errors for Change of Major (50 Replications).
Note. ATT = average treatment on the treated.
For the PSM on selection of program of study, we followed similar procedures. The probit regression included the following independent variables: age, gender, high school rank, first-term credits completed, learning community memberships, honors program member, academic self-efficacy, academic integration, and social integration. The model (Table 7) resulted in a moderately strong goodness of fit (McFadden’s pseudo-R2 = 0.6516). After matching (i.e., nearest neighbor matching algorithm 5 ), the control and treatment groups were statistically balanced. The treatment did not have a significant effect on selecting a program of study (t = 0.179). This nonsignificance was sustained for three other matching methods (stratification, kernel, and radius). Detailed results are shown in Table 8.
Probit Regression for Selection of Program.
Note. n = 827. SE = standard error.
***p < .001.
ATT Estimation With Bootstrapped Standard Errors for Selection of Program (50 Replications).
Note. ATT = average treatment on the treated.
Discussion and Conclusion
Findings of this study demonstrated significant impact of predictive analytics-informed academic advising on change of major. Several interesting findings may need further in-depth discussion. First, this study demonstrated that the use of predictive analytics in academic advising would increase the rate of changing a major among undeclared first-year engineering students. However, changing a major means that these students are leaving engineering major; this is not necessary an undesirable outcome. An alternative way to understand this effect relates to the perspective of students’ overall success instead of the retention rate within engineering department. Multiple previous studies have indicated positive influences of changing a major to students’ overall retention, academic performance, and graduation rate (Foraker, 2012; Kreysa, 2006; Micceri, 2001; Murphy, 2000). Many undeclared students may not doing well after entering the engineering major. However, since they were admitted at the beginning, strong precollege indicators (i.e., ACT scores, high school GPA, etc.) may still make them believe their performance will improve later. The delay of redirection may cost students extra time, effort, and financial input to graduate (Foraker, 2012; Jenkins & Cho, 2012). The predictive analytics enables academic advisors to provide more evidence-based informative feedback based on comparing a particular students’ first-semester performance with similar students from previous years. Such evidence leads students to engage in sensemaking activities (Eliot & Turns, 2011) which in turn guides their decisions toward other majors. Thus, predictive analytics added an objective, data-informed piece of feedback that students are attentive to. In this context, the analytics-based advising was more effective and meaningful than other less recent indicators such as ACT composite score.
Second, although there is an observable difference between treatment and control groups on selecting a program of study (Figure 2), we did not find a statistically significant impact after matching. Previous studies indicated that selection of a program of study at an early stage is associated with better academic achievement in engineering (Kreysa, 2006). The nonsignificant finding in this study can be explained by the followings. First, students in the treatment group had only finished their first term at LMRU by the time of data collection. They may need additional coursework before determining which type of engineering they most enjoy and which field of engineering to enter into. In addition, the LMRU policy did not require undeclared engineering students to decide after their first semester. Second, the predictive analytics were based on a single semester worth of academic performance. With more longitudinal academic performance data from multiple semesters, the predictive analytics may provide a clearer picture of the future performance for an undeclared engineering student.

Trends in the behavior of undergraduate engineering students (2012–2016).
In addition to PSM results, the findings of the logistic regression models are also noteworthy. For example, the logistic regression model of changing a major revealed that academic self-efficacy had the biggest influence among other significant predictors. Undeclared engineering students with higher academic self-efficacy are less likely to change a major. This finding is congruent with many previous studies (Arcidiacono, 2004; Eccles, 1987; Stinebrickner & Strinebrickner, 2011). In our regression model, other significant predictors of changing a major included first-term credits completed and ACT scores. Furthermore, first-term credits completed and academic integration were found significantly and positivity influence students’ likelihood of selecting a program of study. Together, we can conclude that the following four predictors can significantly increase undeclared engineering students’ retention: academic self-efficacy, academic integration, first-term credit completed, and ACT composite scores.
Implications for Policy, Practice, and Future Studies
This study provided important findings regarding the effect of data-driven decision-making on undeclared engineering students. Findings encourage higher education institutions to thoughtfully consider the implementation of data analytics in administration and student service. The following sections summarize implications of promoting data-driven decision-making in higher education as well as recommendations for future studies.
Implications for Policy and Practice
In order to help students achieve successful outcomes such as completing a degree in a timely fashion, data analytics must be considered through institutional policies. Higher education leaders should establish intentional goals when promoting the use of data analytics in student affairs. Especially, this study provided evidence for goal formation (i.e., promoting student success in terms of retention and degree completion). To establish the goals and introduce data analytics into institutional culture, it is very critical to create collaborative environments where successful stories are shared and professional training opportunities are created. Training opportunities are especially important for academic advisors who will directly use the data analytics and interpret the data presented through various sources.
Further, it is also critical for higher education institutions to create and maintain data sharing platforms when implementing data analytics. In this study, the predictive analytics-informed academic advising is made possible only because of the availability of the ample amount of institutional data behind the scenes. Besides academic advisors and other student affairs professionals, higher education researchers also should have access to the raw data, so they can integrate their own research data sets with the analytics data. The access to data sets can be accommodated through robust institutional data warehouse with appropriate levels of access. In this way, researchers can explore various statistical relationships between various meaningful variables. In sum, a strong data governance policy must exist to define the data access policy in a secure and ethical manner that recognizes the needs of higher education leaders, practitioners, researchers, and eventually the needs of promoting student success.
The third implication may relate to future use of the predictive analytics. For example, the usage of predictive analytics can be expanded to create a set of success markers (i.e., specific grades in specific courses) that predict the likelihood of completing an engineering degree in a particular program. If a student does not achieve the predefined grade in a marker course, advisors can reach out to students and make recommendations to address potential problems or encourage the student to select a different engineering program. The successful implementation of this expansion will require further institutional research and professional trainings for academic advisors.
Implications for Future Studies
This study may inspire future studies from the following perspectives. First, findings of this study were based on a short-term (i.e., first semester) outcomes. The longitudinal effect of predictive analytics-informed academic advising should be further examined. A future study may consider tracking students who received predictive analytics-informed academic advising over a longer period of time such as 2, 4, or even a 6-year period. It will be ideal if we could examine the significant effects of predictive analytics on longer term student outcomes such as persistence, completion, and overall GPA.
Second, this study focused on engineering students. However, the data utilized in this study are comparable with data on students from other disciplines. Future studies can use this study as a base to develop various models to examine how predictive analytics may influence academic advising in nonengineering or non-STEM majors. Future studies may also add additional variables that incorporate students’ uniqueness in relation to their specific disciplines or majors. These efforts can contribute to filling the research gap of examining relationships between data analytics and student success.
Last but not the least, this study did find significant effects of introducing analytics-informed academic advising. However, the detailed contexts and processes of advising were not clear (or not controlled). Thus, a qualitative approach may be helpful in further interpret how academic advisors utilized predictive analytics in their advising as well as students’ perspectives of being involved in predictive analytics-informed academic advising.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
