
Editorial
Select search scope: search across all journals or within the current journal

Customer churn prediction is becoming an increasingly important business analytics problem for telecom operators. In order to increase the efficiency of customer retention campaigns, churn prediction models need to be accurate as well as compact and interpretable. Although a myriad of techniques for churn prediction has been examined, there has been little attention for the use of Bayesian Network classifiers. This paper investigates the predictive power of a number of Bayesian Network algorithms, ranging from the Naive Bayes classifier to General Bayesian Network classifiers. Furthermore, a feature selection method based on the concept of the Markov Blanket, which is genuinely related to Bayesian Networks, is tested. The performance of the classifiers is evaluated with both the Area under the Receiver Operating Characteristic Curve and the recently introduced Maximum Profit criterion. The Maximum Profit criterion performs an intelligent optimization by targeting this fraction of the customer base which would maximize the profit generated by a retention campaign. The results of the experiments are rigorously tested and indicate that most of the analyzed techniques have a comparable performance. Some methods, however, are more preferred since they lead to compact networks, which enhances the interpretability and comprehensibility of the churn prediction models.
This paper presents methods for spectral band selection in hyperspectral image (HSI) cubes based on classification of reflectance data acquired from samples of livestock feed materials and ruminant-derived bonemeal. Automated detection of ruminant-derived bonemeal in animal feed is tested as part of an on-going research into development of automated, reliable fast and cost-effective quality control systems. HSI cubes contain spectral reflectance in both spatial dimensions and spectral bands. Support vector machines are used for classification of data in various domains. Selecting a subset of the spectral bands speeds processing and increases accuracy by reducing over-fitting. We developed two methods utilizing divergence values for selecting spectral band sets, 1) evolutionary search method and 2) divergence-based recursive feature elimination approach.
Web usage mining has proven to be an important advance for e-business systems, both by finding web user buying patterns and suggesting ways to improve web user navigation. A primary input for web usage mining is web user sessions that must be constructed from web server logs (called sessionization) when such sessions are not otherwise identified. We use bipartite cardinality matching and a more general integer program to construct sessions. We also propose several variations of our integer program to provide additional insights into session characteristics. For testing, we retrieve 15 months of web server logs and corresponding real sessions from an academic web site. We compare real sessions, results obtained by our optimization models, and results from a commonly-used timeout heuristic. We find our optimization models dominate the timeout heuristic using several comparison measures. Solution time for a typical month is seven hours for our integer program, 30 minutes for our bipartite cardinality matching, and about 1 minute for the heuristic. Although solution time is significantly greater for the integer program, its variations contribute additional analysis of web user behavior.
This paper proposes a framework for automated design of component-based decision tree algorithms. These algorithms are being constructed by interchanging components extracted from decision tree algorithms and their partial improvements. Manual selection of the best-suited algorithm for a specific problem is a complex task because of the huge algorithmic space derived from component-based design. The proposed framework searches through the algorithmic space with an evolutionary algorithm by interchanging components and tuning parameters, and finds a near optimal algorithm for a specific problem. Through experiments we show that using this meta-heuristic is justified in automated component-based algorithm design. This approach is useful not only as an algorithm design help, but also as a technology enhanced learning tool, which aids the understanding of the algorithms.
In statistical research, regression models based on data play a central role; one of these models is the linear regression model. However, this model may give misleading results when data contain outliers. The outliers in linear regression can be resolved in two stages: by using the Mean Shift Outlier Model (MSOM) and by providing a new solution for this model. First, we construct a Tikhonov regularization problem for the MSOM. Then, we treat this problem using convex optimization techniques, specifically conic quadratic programming, permitting the use of interior point methods. We present numerical examples, which reveal very good results, and we conclude with an outlook to future studies.
200 words for Intelligent Data Systems The class imbalance problem is a relatively new challenge that has attracted growing attention from both industry and academia, since it strongly affects classification performance. Research also established that class imbalance is not an issue by itself, but its relationship with class overlapping and noise has an important impact on the prediction performance and stability. This fact has motivated the development of several approaches for classification of imbalanced data (see e.g. [29,39]). In this paper, we present credit card customer churn prediction, an important topic in business analytics, using an ensemble of classifiers. Since this problem is considered as highly imbalanced, we employ different techniques for classification, such as Support Vector Data Description (SVDD) and two-class SVMs. The main idea is to address both class imbalance and class overlapping by stacking different classification approaches, while evaluating the diversity of the individual classifiers considering meta-learning measures. We performed experiments on artificial data sets and one real customer churn prediction problem from a Chilean financial entity, comparing our approach with well-known classification techniques for imbalanced data. The proposed strategy achieves an improvement of 6.1% over the best individual classifier in terms of predictive performance, providing accurate and robust classification models for different levels of balance and noise.
