Abstract
Excessive online shopping cart abandonment rates constitute a major challenge for e-commerce companies and can inhibit their success within their competitive environment. Simultaneously, the emergence of the Internet’s commercial usage results in steadily growing volumes of data about consumers’ online behavior. Thus, data-driven methods are needed to extract valuable knowledge from such big data to automatically identify online shopping cart abandoners. Hence, this contribution analyzes clickstream data of a leading German online retailer comprising 821,048 observations to predict such abandoners by proposing different machine learning approaches. Thereby, we provide methodological insights to gather a comprehensive understanding of the practicability of classification methods in the context of online shopping cart abandonment prediction: our findings indicate that gradient boosting with regularization outperforms the remaining models yielding an F1-Score of 0.8569 and an AUC value of 0.8182. Nevertheless, as gradient boosting tends to be computationally infeasible, a decision tree or boosted logistic regression may be suitable alternatives, balancing the trade-off between model complexity and prediction accuracy.
Keywords
Introduction
To strengthen a company’s position within its competitive environment, marketers need to be able to precisely predict potential customers regarding their purchase and, further, non-purchase behavior. Considering this in the context of online shopping environment, customers frequently place items in their virtual shopping cart for reasons other than immediate purchase. This phenomenon is known as shopping cart abandonment and is particularly apparent in the context of e-commerce; it is the behavioral outcome of consumers placing item(s) in their online shopping cart without making a purchase by completing the checkout process during that online session (Huang et al., 2018; Kukar-Kinney & Close, 2010). Extant literature investigated the behavioral perspective of online shopping cart abandonment by identifying inhibitors to the purchase process: financial risks and concerns about delivery and return policies (Kukar-Kinney & Close, 2010), the usage of shopping carts as organization tools or for entertainment purposes (Kukar-Kinney & Close, 2010), and inhibitors at the checkout stage like perceived transaction inconvenience and privacy intrusion (Rajamma et al., 2009) are—inter alia—the main factors leading to online shopping cart abandonment.
With the spread of the Internet’s commercial usage, the ability to track consumers’ online activities allows companies to collect unbiased information about consumers’ behavior. The detailed records of past usage behaviors comprised by log files and resulting clickstream data can be analyzed by marketers to gain valuable insights. In this context, clickstream data have frequently been modeled to derive implications for website design or advertising efforts (see, for example, Chatterjee et al., 2003 and Montgomery et al., 2004) and further, to predict consumers’ future behaviors, e. g., regarding purchase (see, for example, Bucklin & Sismeiro, 2003 and Moe & Fader, 2004a).
Thus, the antecedents of online shopping cart abandonment are well understood by behavioral literature and clickstream data have been studied by methodological research to analyze consumers’ behavior. The rise of the Internet and the era of big data resulted in an excessive “datafication” (Kelly & Noonan, 2017; Lycett, 2013) of the organizational environment yielding the field of business intelligence comprising data analytics and predictive analytics approaches (Chen et al., 2012). However, despite the richness of clickstream data, prior shopping cart abandonment literature still lacks data-driven methods based on machine learning which make use of this information source to predict such abandoning customers. This might be due to the insufficient awareness of suitable intelligent approaches to extract knowledge from the steadily growing volumes of data (Fayyad et al., 1996).
To address this research gap, we utilize clickstream data of a leading German online retailer to train and subsequently compare different machine learning approaches for the prediction of online shopping cart abandonment (i.e., tree-based methods [more specifically, adaptive boosting, boosted logistic regression, decision tree, gradient boosting with regularization, gradient boosting, gradient boosting with dropout, random forest and stochastic gradient boosting], k-nearest neighbor, naïve bayes, multi-layer perceptron with dropout, and a support vector machine with radial basis kernel). We successfully implement these machine learning methods for online shopping cart abandonment prediction and compare them with logistic regression as a standard non-machine learning benchmark model regarding their predictive performance.
Our article makes several key contributions to the preceding literature. By combining the research fields of both shopping cart abandonment as well as clickstream data analysis with machine learning approaches, we particularly shed light on the practicability of machine learning methods in this application context, as this was neglected by prior research. Furthermore, we provide insights into the characteristics of customers abandoning their shopping cart based on clickstream data that is unsusceptible to self-selection, relatively unobtrusive, and easy to gather. We extensively review literature on classification methods to identify shopping cart abandonments and present validation procedures as well as performance metrics for such methods. Our findings can be useful both for marketing intelligence research by extending the field of machine learning applications in marketing contexts through automatically predicting online shopping cart abandoners and for practitioners to actively prevent such abandonments by several real-time reactions, for example, providing real-time purchase incentives and, moreover, to gain insights into machine learning methods.
The remainder of this article is organized as follows: the subsequent section describes the related work on online shopping cart abandonment and clickstream data. Furthermore, section “Machine learning approaches for classification” summarizes the background on machine learning approaches for classification. Section “Methodology” outlines the methodology comprising a preliminary data analysis and the research design. In sections “Findings” and “Discussion,” we present the findings and discuss both theoretical and practical implications, limitations, as well as directions for future research. Finally, section “Conclusion” draws a conclusion.
Related work
Online shopping cart abandonment
The online shopping cart abandonment phenomenon causes substantial losses of turnover for online retailers (Huang et al., 2018; Rajamma et al., 2009) resulting in a weakened position within their competitive environment. Therefore, extant marketing literature addressed this problem by drawing on a behavioral perspective to identify and understand essential determinants of online shopping cart abandonment; Rajamma et al. (2009) focused on potential inhibitors at the checkout stage and found increased perceived transaction inconvenience (e.g., long registration forms) and high-perceived risk (e.g., perceived security of information asked) to enhance online shopping cart abandonment. Partially, these findings seem to be applicable to new customers which are unfamiliar with the checkout process. Similarly, findings of Kukar-Kinney and Close (2010) indicate that privacy intrusion and security concerns rather lead to the consumers’ decision to buy the product from a stationary offline store. Furthermore, they found the entertainment value of shopping carts, the use of shopping carts as an organization tool, the wait for sale, and the concerns about costs to be antecedents of shopping cart abandonment (Kukar-Kinney & Close, 2010). Their identified determinants were supported by Close and Kukar-Kinney (2010) proving that customers’ tendencies to add items to the online shopping cart for reasons other than immediate purchase are—inter alia—due to organizational purposes. Huang et al. (2018) focused on mobile shopping cart abandonment in their study. They found intrapersonal (i.e., conflicts regarding mobile shopping attributes and low self-efficacy regarding mobile shopping) and interpersonal (i.e., discrepancies from the other’s attitudes to self-attitudes) conflicts to disturb consumers’ emotions during mobile shopping and, in turn, implying shopping cart abandonment. Overall, their findings indicate that the utilized device for online shopping might impact purchase behavior as well. Cho et al. (2006) proved that consumers’ confusion by information overload, high-value consciousness, negative past experiences, intention to conduct price comparisons, and unreliable websites are likely to trigger online shopping cart abandonment. 1
Clickstream data
Drawing on a more holistic perspective of online shopping behavior, further literature shifted away from explanatory behavioral approaches to data-driven methods predicting online purchase behavior in general. Typically, such predictions are based on clickstream data (see, for example, Moe & Fader, 2004a; Sismeiro & Bucklin, 2004; Van den Poel & Buckinx, 2005). Clickstream data model the navigation path a customer takes through the online shop (Montgomery, 2001; Montgomery et al., 2004) and can be extracted from log files which register all requests and information transferred between the customer’s computer and the company’s commercial web server (Bucklin & Sismeiro, 2003).
Examples for using clickstream data to predict online shopping behavior are—inter alia—Moe and Fader (2004a) who proposed a conversion model predicting each customer’s probability of making a purchase based on purchase and visit history. The same authors (Moe & Fader, 2004b) also developed a model for evolving visiting behavior and further, they examined the relationship between visiting frequency and purchasing propensity. They found consumers visiting an e-commerce site more frequently to have a greater propensity to buy (Moe & Fader, 2004b). Van den Poel and Buckinx (2005) predicted purchase behavior and investigated the contribution of different variables: they proved (1) general clickstream variables (i.e., number of days since last visit, and speed of clickstream behavior during last visit), (2) more detailed clickstream variables (i.e., number of accessories [and personal pages and products, respectively] viewed during last visit), (3) demographic variables (i.e., gender and the fact of supplying personal information), and (4) historical purchase behavior (i.e., number of days since last purchase and number of past purchases) to be meaningful predictors. Montgomery et al. (2004) set up different models to predict purchase conversion probability by modeling path information.
Moreover, clickstream data were frequently utilized by research to predict not only purchase behavior but further similar outcome variables. For instance, Bucklin and Sismeiro (2003) investigated drivers affecting the length of time spent viewing a website and the visitor’s decision to continue browsing or to exit the website. Sismeiro and Bucklin (2004) decomposed the purchase process into sequences that must be completed for a purchase to take place (i.e., completion of product configuration, input of personal information, and order confirmation with provision of credit card data) and predicted the probability of completion for each task with covariates of browsing behavior, repeat visitation, use of decision aids, input effort, and information gathering.
Machine learning approaches for classification
Overall, e-commerce as a research subject is suitable for the application of machine learning approaches as proposed by Kohavi and Provost (2001): online retailers can easily and inexpensively collect rich data with respect to the online behavior of customers (i.e., clickstream data) and, further, implement data mining and machine learning applications since political and social barriers are substantially lower than for traditional businesses. Consequently, typical problems for successfully applying machine learning (i.e., the need for a large volume of controlled and reliable data, data with sufficient descriptions, the ability to evaluate results, and to integrate applications successfully) are reduced by the characteristics of e-commerce environment (Kohavi & Provost, 2001).
Machine learning constitutes a new paradigm within data science research and emerged in the course of the artifical intelligence era, which, in turn, was first coined by Samuel (1959) describing it as “the programming of a digital computer to behave it in a way which, if done by human beings [. . .], would be described as involving the process of learning”. In this context, learning may be understood as the automatic search for more useful representations of data regarding a specific task (Chollet & Allaire, 2018). Machine learning algorithms and systems are consequently trained rather than explicitly programmed. During this process, these systems find statistical structure in given examples which are relevant to the task and derive rules for automating the task using guidance from a feedback signal (Cui et al., 2006). Thereby, classification algorithms are types of supervised learning approaches within machine learning, which predict a qualitative response for an observation, that is, they assign an observation to a category (James et al., 2013): Formally, let
Drawing on the online shopping cart abandonment problem, the prediction of purchasers and non-purchasers (i.e., customers abandoning their shopping cart) can be considered a binary classification task. Common machine learning approaches for binary classification include—inter alia—tree-based methods, support vector machines, naïve bayes, k-nearest neighbor, and neural networks. The approaches are explained in detail hereinafter.
Tree-based approaches
One of the most common machine learning approaches are tree-based methods which descend from single decision trees, as proposed by Breiman et al. (1984). Basically, decision trees are flowchart-like structures that generate “if-else” rules and thereby allow for prediction of observation classes. Thereby, classification and regression tree models follow a recursive top-down approach in which binary trees aim to partition the predictor space with predictor variables x1, . . . , xk into subsets in which the distribution of the dependent variable
Generally, single decision trees have the advantage of being easy to interpret and to understand (Moro et al., 2014). However, they frequently lead to overfitting, that is, the model learns to identify specific characteristics of the training data which are irrelevant or even obstructive for the classification of unknown data (Friedman, 2001; Srivastava et al., 2014). This results in drawbacks of predictive performance and less expressiveness of the models. Ensemble learning methods that construct several individually trained decision trees and combine their results into a classifier outperforming the single predictions (Opitz & Maclin, 1999; Rokach, 2010) may offer a solution to this problem. In this context, two widely used methods of aggregating trees are boosting and bagging.
In boosting, a family of algorithms converts weak learners (i.e., models that achieve accuracy just above random guessing) to strong learners with a powerful predictive capacity. The idea is to train weak learners sequentially with each weak learner trying to correct its predecessor (Schapire et al., 1998). Thus, each decision tree is built using feedback from previously grown trees (James et al., 2013). Popular boosting algorithms include adaptive boosting “AdaBoost” (Freund & Schapire, 1997), boosted logistic regression “LogitBoost” (Friedman et al., 2000), gradient boosting machines “GB” (Friedman, 2001, 2002), and stochastic gradient boosting “SGB” (Friedman, 2002). 2 For instance, AdaBoost as a basic boosting algorithm makes predictions by combining the output of weak learners to a weighted sum and putting higher weights on incorrectly classified instances
with the weak hypothesis
In contrast to boosting, bagging (or bootstrap aggregating) grows successive trees independently from earlier trees, that is, each tree is constructed using a bootstrap sample of the data and, hence, a majority vote is taken for prediction (Breiman, 1996). Random forests add an additional layer of randomness to bagging and change how the trees are constructed; in standard decision trees, each node is split using the best split among all predictor variables, whereas, in random forests, the nodes are split using the best among a subset of predictors randomly chosen at that node (Breiman, 2001; Liaw & Wiener, 2002). Due to the recursive structure of tree-based methods, they often capture interaction effects between variables. However, since we focus on the performance of the models and not the importance of specific variables, we will not consider interaction effects further in our study.
Overall, tree-based methods have been found to outperform other established approaches across a variety of different classification tasks such as Internet protocol (IP) traffic flow classification (Williams et al., 2006), customer churn prediction (Vafeiadis et al., 2015), or—similar to our context—prediction of online purchase intention (Bogina et al., 2019; Boroujerdi et al., 2014; Zheng & Liu, 2018). They are particularly favorable since ensemble methods are able to reduce both bias and variance of the single learning algorithms: While individual models may get stuck in local minima, a weighted combination of several different local minima—produced by ensemble methods—are able to minimize the risk of choosing the wrong local minimum (Dietterich, 2002).
Support vector machines
Aside from tree-based methods, support vector machines are powerful tools for classification tasks (James et al., 2013). The basic support vector machine is solving pattern recognition problems by mapping data into a multidimensional input space and constructing an optimal hyperplane that separates the space into homogeneous partitions 3 (Cortes & Vapnik, 1995; Vapnik, 1982). Predictions of new instances are then classified into those partitions. The support vector machine aims at constructing a classifier in the form of
where
We used a support vector machine with radial basis kernel for the comparison of machine learning models. However, support vector machines may become computationally infeasible on very large datasets like clickstream data (L’Heureux et al., 2017).
Naїve Bayes
The naїve bayes approach is a basic classifier based on applying the Bayes’ theorem with the naїve assumption that the attributes are conditionally independent (Duda et al., 1973). The classifier assigns a new case to a class label
Naïve bayes as a generative classifier is frequently utilized for classification tasks due to its simplicity, efficiency, and efficacy (Muhammad & Yan, 2015).
K-nearest neighbor
Another basic approach, the k-nearest neighbor algorithm, classifies an observation by a majority vote of the observation’s neighbors (Cover & Hart, 1967). The underlying assumption of the algorithm is that observations which lay closely together within the predictor space (i.e., neighbors) will have the same class label. Thus, the classifier weights the class of the nearest neighbors strikingly high to predict the class label of an unclassified sample (Cover & Hart, 1967). The class is thereby assigned by taking the majority vote of the k nearest neighbors, with k being the number of neighbors that are considered during the classification task. The nearest neighbors are determined with the help of arbitrary distance functions (e.g., Euclidian distance
and
K-nearest neighbor as a local learning approach may be suitable for online shopping cart abandonment prediction tasks since it is able to alleviate the challenge of imbalanced data (L’Heureux et al., 2017).
Artificial neural networks
Artificial neural networks are highly parallelized computer systems comprising process units (i.e., neurons) located on process layers with numerous weighted interconnections performing a learning process to create meaningful data representations (Jain et al., 1996). Regarding the concept of deep learning, artificial neural networks may use a number of hidden process layers (the depth of a network) between input and output layer containing non-linear operations in hierarchical architectures to learn characteristics and recognize patterns from given data (Bengio, 2009; Deng, 2011; Hinton et al., 2006). The concept of learning within deep learning (or artificial neural networks, respectively) describes a process of updating the network architecture and the weights of the neuron connections (Jain et al., 1996). To improve the performance, the optimizer is implementing a backpropagation algorithm to minimize the discrepancy between the actual and the target output vector (i.e., the loss score) by adjusting the weights (Rumelhart et al., 1986; Schmidhuber, 2015). To avoid overfitting, a regularization method called dropout can be integrated in the network which randomly sets a share of its output per layer to zero (Srivastava et al., 2014).
Concerning their connection structure (i.e., topology), neural network architectures can be distinguished between feedforward networks (e.g., multi-layer perceptrons; Deng, 2011; Q. Zhang et al., 2018) with neuron connections running to the output layer acyclically and recurrent networks (e.g., long short-term memories; Hochreiter & Schmidhuber, 1997) containing backward connections to build cyclic architectures (Jain et al., 1996; Schmidhuber, 2015). The most commonly used feedforward neural networks—multi-layer perceptrons—can be defined as
where
Multi-layer perceptrons were found to outperform other machine learning approaches for purchase intention prediction only after balancing the class distribution with oversampling (Sakar et al., 2019), since deep learning approaches are frequently sensitive to class imbalance (L’Heureux et al., 2017).
Methodology
Preprocessing and preliminary data analysis
The purpose of this study is to predict shopping cart abandonment by making use of machine learning. The machine learning models explained in section “Machine learning approaches for classification” are compared to find the best classifier for this task. The clickstream data were gathered from server log files of a leading German online retailer, which primarily distributes fashion. The data were created by the online retailer through extracting the customers’ chronological online shop activities out of sequential log files. Each log file observation comprised one action or activity (e.g., a click) of a certain customer such as adding a product to the cart or clicking on a product to view its details. Subsequently, each customer’s activities during a session were assigned to summarizing variables. Hence, all activities of a customer were aggregated to one observation with different variables describing the session. Thereby, a session is a period of sustained web browsing or a sequence of the user’s page viewings until the user exits the online shop (Montgomery et al., 2004). The data comprise 3,511,037 observations or sessions between 1 February 2019 and 30 April 2019, that is, 3 months. Furthermore, the data contain 18 explanatory variables for each observation or session listed in Table 1 many of which are consistent with Van den Poel and Buckinx (2005) findings. We are only interested in visitors who made use of the virtual shopping cart during the session, that is, who placed item(s) in their cart. In line with Close and Kukar-Kinney (2010), shopping cart usage is thus defined as necessary precondition for shopping cart abandonment. Thus, we filtered out customers who did not add any items to their shopping cart during the session, so-called just-browsing customers, and 821,048 observations (23.38%) remained. We modeled the dependent variable—shopping cart abandonment—as a dummy variable using the information about the customer’s compiled and ordered shopping carts (variables BASKETS_BB and BASKETS) during the session
Variables of clickstream data.
BASKETS: Number of Carts Compiled; LOGS: Number of Logins; LOGS_CUST_STEP2: Number of Existing Customers’ Logins to the Second Step of the Ordering Process; LOGS_NEWCUST_STEP2: Number of New Customers’ Logins to the Second Step of the Ordering Process; MOBILE_CUST: Customer Accessing through Mobile Phone; NEW_CUST: New Customer; PIS: Number of Overall Page Viewings; PIS_AP: Number of Shopping Cart Page Viewings; PIS_DV: Number of Detailed Product Page Viewings; PIS_PL: Number of Category Overview Page Viewings; PIS_SDV: Number of Detailed Product Page Viewings Using Search Function; PIS_SHOPS: Number of Department Page Viewings; PIS_SR: Number of Search Results Page Viewings; POSITIONS: Number of Product Types; QUANTITY: Number of Items; WEB_CUST: Customer Accessing through Desktop.
Our data contain 520,653 (63.41%) observations of shopping cart abandonments (or non-purchasers, respectively) and 300,395 (36.59%) observations of purchasers. Hence, the dataset is relatively balanced. We excluded the variable for the number of ordered shopping carts (BASKETS_BB) and the value of ordered shopping carts (VALUE_BB) further for prediction. 4
Figure 1 illustrates the relationship between the page viewing and login variables by demonstrating the customer’s clickstream in the online shop; the customer typically starts browsing departments (PIS_SHOPS), then selects a certain category within a department (PIS_PL), and further, chooses a certain product within a category (PIS_DV). Optionally, the customer uses the shop’s search engine (PIS_SR) to look systematically for a specific product (PIS_SDV). To make a purchase, the customer can either directly sign in (LOGS) or check the items in the shopping cart (PIS_AP) first and then sign in and hence, proceed to the second step of the purchasing process (LOGS_CUST_STEP2 or LOGS_NEWCUST_STEP2). However, signing in to the second step of the purchasing process does not necessarily lead to a purchase of the customer.

Main clickstream of customers in the online shop.
Nevertheless, with respect to the descriptive statistics in Table 2, we find that existing customers (or new customers, respectively) which subsequently make a purchase sign in to the second step of the ordering process approximately 5.93 times (or 4.46 times, respectively) more often than non-purchasers. Generally, purchasers sign in more often (1.03 logins on average) than non-purchasers (0.93 logins on average). This might indicate that the cause for shopping cart abandonment frequently occurs before the customer proceeds to the checkout stage.
Descriptive statistics of clickstream data.
BASKETS: Number of Carts Compiled; LOGS: Number of Logins; LOGS_CUST_STEP2: Number of Existing Customers’ Logins to the Second Step of the Ordering Process; LOGS_NEWCUST_STEP2: Number of New Customers’ Logins to the Second Step of the Ordering Process; MOBILE_CUST: Customer Accessing through Mobile Phone; NEW_CUST: New Customer; PIS: Number of Overall Page Viewings; PIS_AP: Number of Shopping Cart Page Viewings; PIS_DV: Number of Detailed Product Page Viewings; PIS_PL: Number of Category Overview Page Viewings; PIS_SDV: Number of Detailed Product Page Viewings Using Search Function; PIS_SHOPS: Number of Department Page Viewings; PIS_SR: Number of Search Results Page Viewings; POSITIONS: Number of Product Types; QUANTITY: Number of Items; SD: standard deviation; WEB_CUST: Customer Accessing through Desktop.
Furthermore, the number of purchasers’ overall page viewings is 2.09 times higher than of non-purchasers on average. Overall, customers abandoning their shopping cart browse less pages than purchasers—regardless of the pages’ type. Particularly, the median reveals that there are significant differences regarding the number of page viewings between purchasers and abandoners: the median of abandoners’ overall page viewings is 12, 1 for department viewings, and 0 for all other types of page viewings. In contrast, purchasers’ median for overall page viewings is 35, 6 for department viewings, and for example, 2 for shopping cart viewings.
On average, purchasers add more items and different product types (3.48 and 3.38, respectively) to their shopping cart than non-purchasers (2.95 and 2.88, respectively).
There is a larger absolute (48,839) and relative (9.38%) proportion of new customers among the observations of shopping cart abandonments than among those making a purchase (15,387 observations or 5.12%, respectively). Moreover, there is a larger proportion of mobile shoppers among customers abandoning their shopping cart (45.85%) compared to the observations of purchasers (28.1%). The latter descriptive findings are consistent with the results of preceding (behavioral) research: for example, as argued earlier, Huang et al. (2018) proved that online shopping cart abandonment occurs more frequently for customers using a mobile device due to high emotional ambivalence. Moe and Fader (2004a) found that—among new customers—online conversion rate is lower as purchasing thresholds and perceived risks are high for unexperienced visitors.
Experimental setup
Since each machine learning approach and its subsequent refinements and modifications exhibit individual strengths and weaknesses in dependence of the underlying data and the requested task, it is highly recommended in the machine learning literature to compare and test different algorithms (Moro et al., 2014; Razi & Athappilly, 2005). Thus, we compared different models of those proposed in section “Machine learning approaches for classification” to predict shopping cart abandonment for our data, listed in Table 3. In addition, we included a standard logistic regression model in our comparison serving as a non-machine learning benchmark method.
Machine learning approaches for comparison.
To estimate and, hence, validate the models, we randomly partitioned the data into a training and a test subset in a 67/33 ratio, that is, 67% (or 550,098 observations, respectively) of the data are used as training data and 33% (or 270,950 observations, respectively) are used as test data.
We performed
Furthermore, to validate and evaluate our models’ performance, we considered different performance metrics that indicate the models’ predictive ability. In a binary decision problem, the classifier labels observations as either positive or negative. Consequently, the classification procedure yields four different outputs in a
However, recent research shifted away from solely presenting accuracy results since accuracy assumes balanced class distribution and equal error costs (i.e., Type I errors are equivalent to Type II errors) which is rarely the case in real world applications (Davis & Goadrich, 2006; Provost & Fawcett, 1997). To address these problems, a receiver operating characteristics (ROC) curve and thus, the area under the ROC curve (AUC)
5
has been increasingly used by the machine learning community since they are insensitive to changes in class distributions and scale-invariant (Bradley, 1997; Fawcett, 2006). A ROC graph is a two-dimensional depiction of classification performance to measure different classifiers’ performances and captures the trade-off between benefits (i.e., true positives) and costs (i.e., false positives) (Fawcett, 2006). It is created by plotting the true positive rate (TPR) (or sensitivity or recall, respectively) against the false positive rate (FPR) (or
The classifier’s AUC value is a portion of the area of the unit square and its value ranges from 0.0 to 1.0 (perfect classification). It should be higher than 0.5 which equals the AUC of an uninformative classifier (Bradley, 1997; Fawcett, 2006). An important statistical property of the AUC is that a classifier’s AUC is equivalent to the probability that the classifier will rank a randomly chosen positive observation higher than a randomly chosen negative observation (Fawcett, 2006).
An alternate performance measure is the
Ideally, the performance measure is chosen by properly reflecting the investigation’s aims to avoid misleading conclusions. Since our data are relatively balanced, it seems reasonable to consider accuracy as a basic performance metric. However, as we intend to convert customers abandoning their shopping carts into purchasers our main aim is to correctly classify actual positives (i.e., observations of shopping cart abandonments) by minimizing the Type I error. Consequently, the higher the recall the less false negatives (i.e., shopping cart abandonments classified as purchasers) have been predicted. Besides, we intend to maximize the proportion of actual positives among the predicted positives by minimizing the Type II error, that is, purchasing customers should not be classified as non-purchasers. Thus, the higher the precision the less false positives have been predicted. The
Although prediction accuracy (i.e., AUC, F1-Score, or accuracy) is frequently the main decision criterion when comparing different machine learning models, the models’ complexity in terms of computation time and computation effort (e.g., numbers of hyperparameters to be optimized) is of similar importance regarding the application in practice and should therefore be considered as well (Doshi-Velez & Kim, 2017; Guidotti et al., 2019; Tambe et al., 2019).
Findings
Drawing on the training results in Table 4, gradient boosting with regularization outperformed the remaining approaches with an AUC of 0.9008. The final gradient boosting model’s fitted hyperparameters did not include the lasso regression technique (L1 regularization) but made use of the ridge regression technique (L2 regularization). The gradient boosting with tree base learners and random forest yielded comparable results (AUC of 0.8953 and 0.8954, respectively) whereas naïve bayes and boosted logistic regression realized the lowest AUC values (0.8218 and 0.8381, respectively).
Training data results.
The highest AUC value is marked in bold. AdaBoost: Adaptive Boosting; DT: Decision Tree; GBDropout: Gradient Boosting with Dropout; GBReg: Gradient Boosting with L1 and L2 Regularization; GBTree: Gradient Boosting with Tree Base Learners; KNN: k-Nearest Neighbor; LogitBoost: Boosted Logistic Regression; MLPDropout: Multi-Layer Perceptron Network with Dropout; NB: Naïve Bayes; RF: Random Forest; SGB: Stochastic Gradient Boosting; SVMRadial: Support Vector Machine with Radial Basis Kernel.
With 40 GB RAM.
Regarding estimation time, the benchmark logistic regression, decision tree and boosted logistic regression performed the fastest 10-fold cross validation to optimize the hyperparameters (20.3, 225.07, and 380.0 s, respectively). The support vector machine and adaptive boosting were the most time-consuming models to estimate (1,306,838.6 and 703,903.9 s, respectively). Gradient boosting with regularization yielded a moderate estimation time (4,021.28 s) and thus, provides an appropriate trade-off between AUC and estimation time.
Since we are rather interested in the fitted models’ performances on new and unknown data, the test data result in Table 5 exhibit a higher practical relevance than the preceding results: similarly to the training data results, the gradient boosting model with regularization was superior to the remaining models regarding the test data. It yielded the best AUC (0.8182) and accuracy (82.29%) results. In line with these findings, the
Test data results.
For each column, the highest value is marked in bold. AdaBoost: Adaptive Boosting; DT: Decision Tree; GBDropout: Gradient Boosting with Dropout; GBReg: Gradient Boosting with L1 and L2 Regularization; GBTree: Gradient Boosting with Tree Base Learners; KNN: k-Nearest Neighbor; LogitBoost: Boosted Logistic Regression; MLPDropout: Multi-Layer Perceptron Network with Dropout; NB: Naïve Bayes; RF: Random Forest; SGB: Stochastic Gradient Boosting; SVMRadial: Support Vector Machine with Radial Basis Kernel.
Although naïve bayes realized an extremely high recall (0.9996), its precision (0.6351) is just slightly better than random guessing. This is due to its negligible Type I error (i.e., 68 abandonments classified as purchasers [0.0004% of all abandonments]) and its substantial Type II error (i.e., 98,677 purchasers classified as abandonments [99.52% of all purchasers]). Consequently, by focusing exclusively either on precision or recall, one could draw misleading conclusions regarding model selection. The
Similarly, albeit the decision tree classified a high proportion of purchasers correctly and only 12,688 (i.e., 12.80% of all purchasers) wrong, it categorized 55,634 cart abandonments as purchasers (i.e., 32.38% of all abandonments). Thus, due to its high Type I error, its recall is extremely low (0.6762), but it realized the highest precision value of all models (0.9015).
Generally, our results indicate a substantial predictive ability of the most tree-based methods (i.e., gradient boosting with regularization [and linear base learners], gradient boosting [with tree base learners], gradient boosting with dropout [and tree base learners], and random forest) compared with the remaining machine learning approaches. The latter were outperformed by tree-based models with regard to all relevant performance metrics (AUC, accuracy, and
Logistic regression as a non-machine learning benchmark approach yielded the lowest
Moreover, the k-nearest neighbor algorithm as a basic machine learning approach outperformed more sophisticated algorithms like the multi-layer perceptron, the stochastic gradient boosting, and adaptive boosting with respect to its AUC value (0.7962).
Discussion
Our findings contribute to a deeper understanding regarding the successful implementation of machine learning methods for predicting online shopping cart abandoners with a strong forecast performance to apply marketing techniques in real-time to convert them to purchasers. Thus, we discuss our findings’ theoretical contribution and practical implications in this section. We also discuss limitations and propose suggestions for future research.
Theoretical contribution
Overall, we fill a research gap by identifying suitable machine learning approaches for online shopping cart abandonment prediction not only in terms of accuracy but, further, in terms of practicability. Thereby, we contribute to literature in several ways. First, we are able to characterize customers abandoning their shopping cart descriptively with our data. Preceding literature on shopping cart abandonment (e.g., Close & Kukar-Kinney, 2010; Huang et al., 2018; Kukar-Kinney & Close, 2010) primarily shed light on behavioral aspects of the abandonment process with experimental designs. In contrast, our research deals with unbiased clickstream data comprising an exceptionally high number of observations. Our data indicate that there is a higher proportion of new customers and mobile shoppers among customers abandoning their shopping carts compared to purchasers, whereas the latter add more items to their shopping cart and view an increased number of pages on average.
Second, we contribute to literature by proposing a broad range of machine learning models to compare their performance regarding online shopping cart abandonment prediction and, thus, to predict future customers abandoning their shopping carts in real-time. Prior literature either drew on a behavioral perspective to understand the antecedents of shopping cart abandonment or predicted—more generally—purchase behaviors with conservative approaches and less observations (see, for example, Huang et al., 2018; Kukar-Kinney & Close, 2010; or Sismeiro & Bucklin, 2004). For our data, the gradient boosting with regularization yielded the highest accuracy (82.29%). However, with respect to our main aim, to minimize the Type I error (i.e., abandoners falsely classified as purchasers) and the Type II error (i.e., purchasers falsely classified as abandoners), we focused on the
Overall, we found tree-based methods to be superior to the remaining machine learning approaches and logistic regression as a benchmark non-machine learning approach aligning with prior research comparing machine learning approaches in different application fields like customer churn prediction or phishing detection (Abu-Nimeh et al., 2007; Caruana & Niculescu-Mizil, 2006; Vafeiadis et al., 2015) and—similar to our context—prediction of online purchase intention (Bogina et al., 2019; Boroujerdi et al., 2014; Zheng & Liu, 2018). Thus, we complement the literature on machine learning comparisons in a marketing context.
Moreover, despite the striking importance of prediction accuracy as a decision criterion for appropriate machine learning approaches, the models’ practicability with respect to modeling complexity as an essential criterion is of particular importance (Doshi-Velez & Kim, 2017; Guidotti et al., 2019; Tambe et al., 2019) but, at the same time, is often neglected by current research. Thus, we considered the models’ complexity in terms of computation time and computation effort (e.g., numbers of hyperparameters to optimize) to add to literature. Thereby, the decision tree approach and boosted logistic regression yielded only slightly worse AUC results compared to gradient boosting with regularization and, simultaneously, their complexity in terms of both computation effort and time was rather low. Hence, in case of online shopping cart abandonment prediction, a decision tree model and boosted logistic regression perform well in balancing the trade-off between accuracy and complexity. Furthermore, as stated by prior literature, we found the support vector machine approach to be extremely computationally infeasible (L’Heureux et al., 2017) despite its acceptable prediction accuracy.
Practical implications
Our research may help to gather a comprehensive understanding of machine learning approaches for prediction or classification, particularly with regard to online shopping cart abandonment prediction. More specifically, our research provides multifold practical implications for decision makers.
Since research about advanced machine learning approaches in marketing contexts is still in its infancy (e.g., Cheung et al., 2003 and Cui et al., 2006), we reviewed relevant literature to provide an introduction to such models, its potential applications, as well as performance metrics, and common methods for validation: for machine learning models,
Aside from pointing out methodological aspects, we drew on an economical perspective to enhance an organization’s turnover; with regard to our data, the mean value of purchasers’ ordered shopping carts (VALUE_BB) is €271.73 and they add 3.479 items into their shopping cart on average and thus, we expect the online retailer’s sales loss for each shopping cart abandonment to be around €230 with 2.945 items in their shopping cart on average. Therefore, we determined a suitable approach to correctly identify shopping cart abandonments as well as purchasers; our findings indicate that gradient boosting with regularization outperformed the remaining approaches. Organizations can implement this method to predict non-purchasers in real-time when a sufficient amount of information about the customer’s activities during the session has been collected. Overall, we found particularly tree-based machine learning approaches such as random forest or gradient boosting to outperform traditional classification approaches such as logistic regression and decision tree, which are frequently utilized by practitioners.
Drawing on an overall practicability perspective, decision makers may take a slight loss in prediction accuracy into account if, instead, the model’s complexity in terms of computation time and effort is substantially lower: in our application context, decision tree and boosted logistic regression yielded acceptable prediction results and their computation effort was substantially lower compared to gradient boosting methods.
Limitations and future research
Our research is subject to limitations which stimulate further research. First, the set of useful variables for prediction was limited. With respect to extant literature (see, for example, Bucklin & Sismeiro, 2003; Moe & Fader, 2004a; Van den Poel & Buckinx, 2005), we expect, for example, demographic variables, historical purchase behavior, or the time customers spend on the single pages to be informative variables. Furthermore, we did not have information about the customers’ identity and thus, could not determine whether there were recurring customers. However, this information could be of great interest for analyzing online behavior and predicting shopping cart abandonment. For instance, Huang et al. (2018) anticipated that some customers might use the mobile phone for initial purchase stages (i.e., browsing and collecting information) and then switch to the computer for completing the purchase. However, such customers are listed as two distinct sessions in the current data. Another missing information concerns the value of abandoned shopping carts. While there is a variable that indicates the value of ordered carts (i.e., VALUE_BB), the value of abandoned carts can only be estimated. In line with extant literature on shopping cart abandonment (e.g., Close & Kukar-Kinney, 2010; Kukar-Kinney & Close, 2010), it can be assumed that the value of ordered items influences abandoning rates and, thus, could aid the prediction of such. Moreover, if detailed information about spent time and further, the chronological order of customers’ actions in the online shop would be available, we could decompose the session into sequences or segments. Then, we could determine a critical point in the customer’s session in which abandonment can be predicted reliably with the
Second, we excluded just-browsing customers from our investigation. A possible direction for future research could be to conduct a multi-class classification by differentiating between purchasers, abandonments, and just-browsing customers, similar to the cluster analysis of Moe (2003).
Third, the models’ performance strongly depends on the optimized hyperparameters which may be a time-consuming procedure for some of the models. Therefore, we considered only a limited range of possible hyperparameter values. Moreover, other values of
Finally, a real-time implementation requires a certain amount of data to be collected before the model can make a reliable decision.
By implementing these models, companies may detect shopping cart abandoners in real-time and, subsequently, convert some of them into purchasers by making use of targeted marketing measures such as individual chat pop-ups, coupons or special discounts. For instance, Close and Kukar-Kinney (2010) suggest human-human interactions (i.e., live chats with employees or other online shoppers) to avoid shopping cart abandonment. These could pop-up on the website if the online user is predicted to abandon by the machine learning model. Therefore, future research is recommended to test whether pop-up messages and offers impact customers’ online shopping behavior and can prevent online shopping cart abandonment.
Conclusion
Online shopping cart abandonment can inhibit corporate growth and hence, harm a company’s success within its competitive environment. Simultaneously, the emergence of the Internet’s commercial usage leads to the ability to track consumers’ online activities and online behavior resulting in clickstream data.
Thus, to identify online shopping cart abandoners by extracting valuable knowledge from such clickstream data, we proposed different machine learning approaches. We analyzed data of a German online retailer comprising 821,048 observations and fitted the models using 10-fold cross validation. Thereby, our article contributes to extant literature by combining research fields of both online shopping cart abandonment and clickstream data with machine learning approaches.
Our data indicate that among customers abandoning their shopping carts there is a higher proportion of new customers and mobile shoppers compared to purchasers, whereas the latter add more items to their shopping cart and have a higher number of page viewings on average. Moreover, our comparison results prove that gradient boosting with regularization is a suitable method to distinguish between abandonments and purchasers yielding an AUC of 0.8182, an
Nevertheless, research on clickstream data combined with machine learning approaches is still in its infancy—particularly in a marketing context. Thereby, machine learning will be inevitable for e-commerce businesses to be successful in the long-term and the analysis provided in this article shall stimulate further research on this topic.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Notes
Appendix
confusion matrices
| Model | Prediction | Actual | |
|---|---|---|---|
| 0 (Purchaser) | 1 (Abandonment) | ||
| Logistic regression | 0 (Purchaser) | 83,817 | 41,722 |
| 1 (Abandonment) | 15,335 | 130,076 | |
| AdaBoost | 0 (Purchaser) | 62,009 | 21,005 |
| 1 (Abandonment) | 37,143 | 150,793 | |
| LogitBoost | 0 (Purchaser) | 72,036 | 34,692 |
| 1 (Abandonment) | 27,116 | 137,106 | |
| DT | 0 (Purchaser) | 86,464 | 55,634 |
| 1 (Abandonment) | 12,688 | 116,164 | |
| GBReg | 0 (Purchaser) | 79,385 | 28,209 |
| 1 (Abandonment) | 19,767 | 143,589 | |
| GBTree | 0 (Purchaser) | 77,662 | 27,875 |
| 1 (Abandonment) | 21,490 | 143,923 | |
| GBDropout | 0 (Purchaser) | 78,294 | 28,352 |
| 1 (Abandonment) | 20,858 | 143,446 | |
| KNN | 0 (Purchaser) | 75,687 | 29,383 |
| 1 (Abandonment) | 23,465 | 142,415 | |
| MLPDropout | 0 (Purchaser) | 73,803 | 27,869 |
| 1 (Abandonment) | 25,349 | 143,929 | |
| NB | 0 (Purchaser) | 475 | 68 |
| 1 (Abandonment) | 98,677 | 171,730 | |
| RF | 0 (Purchaser) | 77,903 | 28,197 |
| 1 (Abandonment) | 21,249 | 143,601 | |
| SGB | 0 (Purchaser) | 74,409 | 29,217 |
| 1 (Abandonment) | 24,743 | 142,581 | |
| SVMRadial | 0 (Purchaser) | 72,724 | 24,427 |
| 1 (Abandonment) | 26,428 | 147,371 | |
AdaBoost: Adaptive Boosting; DT: Decision Tree; GBDropout: Gradient Boosting with Dropout; GBReg: Gradient Boosting with L1 and L2 Regularization; GBTree: Gradient Boosting with Tree Base Learners; KNN: k-Nearest Neighbor; LogitBoost: Boosted Logistic Regression; MLPDropout: Multi-Layer Perceptron Network with Dropout; NB: Naïve Bayes; RF: Random Forest; SGB: Stochastic Gradient Boosting; SVMRadial: Support Vector Machine with Radial Basis Kernel.
