
Editorial
Select search scope: search across all journals or within the current journal

Many distance or similarity measures have been proposed for time series similarity search. However, none of these measures is guaranteed to be optimal when used for 1-Nearest Neighbor (NN) classification. In this paper we study the problem of selecting the most appropriate distance measure, given a pool of time series distance measures and a query, so as to perform NN classification of the query. We propose a framework for solving this problem, by identifying, given the query, the distance measure most likely to produce the correct classification result for that query. From this proposed framework, we derive three specific methods, that differ from each other in the way they estimate the probability that a distance measure correctly classifies a query object. In our experiments, our pool of measures consists of Dynamic Time Warping (DTW), Move-Split-Merge (MSM), and Edit distance with Real Penalty (ERP). Based on experimental evaluation with 45 datasets, the best-performing of the three proposed methods provides the best results in terms of classification error rate, compared to the competitors, which include using the Cross Validation method for selecting the distance measure in each dataset, as well as using a single specific distance measure (DTW, MSM, or ERP) across all datasets.
To meet the need of extracting cluster boundary from mixed attribute data in the field of data analysis, we propose a cluster boundary detection algorithm for mixed attribute data sets in this research, named CHASM (Cluster Boundary Detection Algorithm based on Shadowed Set). Based on the structure of clusters, the CHASM defines a new objective function according to the data set which is categorized into three collections, i.e. core, exclusion and shadow. Then CHASM updates the centroid information of clusters based on the variance of contribution degree among these collections to the clusters centroids. Finally, in an iterative optimization process, the CHASM can extract its shadow sets from each cluster to form the boundary of clusters. The experimental results, on both the synthetic data and real data with mixed attributes, numerical attributes and categorical attributes, show that CHASM can effectively detect cluster boundary with higher or similar accuracy to its rival methods. Furthermore, the CHASM can eliminate noise effectively.
In the last 10 years, the information generated on weblog sites has increased exponentially, resulting in a clear need for intelligent approaches to analyse and organise this massive amount of information. In this work, we present a methodology to cluster weblog posts according to the topics discussed therein, which we derive by text analysis. We have called the methodology
In this paper, an iterative self-training Support Vector Machine (SVM) algorithm combined feature re-extraction is proposed for semi-supervised learning, which only needs a small set of labeled samples to train classifier and is thus very useful in Brain-Computer Interface (BCI) design. Two methods, the model selection based self-training and the confidence criterion, respectively, is also proposed for searching the best parameter pair of SVM and selecting the most useful unlabeled data to expand the labeled training data set. The Dataset IVa of BCI Competition III, is presented to demonstrate the validity of our algorithm with statistical significance test. As an iterative algorithm, experimental results of the proposed algorithm show the validity of re-extracting feature and the robustness of the feature to the noise. In addition, the convergence of the proposed algorithm and the validity of the method measuring the consistency of the feature are also demonstrated in experiments.
Variable selection is crucial for improving interpretation quality and forecasting accuracy. To this end, it is very interesting to choose an effective dimension reduction technique suitable for processing data according to their specificity and characteristics. In this paper, the problem of variable selection for linear and nonlinear regression is deeply investigated. The curse of dimensionality issue is also addressed. An intensive comparative study is performed between Support Vector Regression(SVR) and Random Forests (RF) for the purpose of variable importance assessment then for variable selection. The main contribution of this work is twofold: to expose some experimental insights about the efficiency of variable ranking and selection based on SVR and on RF, and to provide a benchmark study that helps researchers to choose the appropriate method for their data. Experiments on simulated and real-world datasets have been carried out. Results show that the SVR score ∂ Gα is recommended for variable ranking in linear situations whereas the RF score is preferable in nonlinear cases. Moreover, we found that RF models are more efficient for selecting variables especially when used with an external score of importance.
Association rule mining meeting a variety of measures is regarded as a multi-objective optimization problem rather than a single objective optimization problem. The convergent speed of traditional multi-objective algorithms such as genetic algorithm is slow and the efficiency of these algorithms is low. Furthermore, the rules generated by traditional multi-objective algorithms are too large to be efficiently analyzed and explored in any further process. Bat algorithm is a new efficient global optimal algorithm whose convergence is superior to binary particle swarm optimization (BPSO) and genetic algorithm. This paper discusses the application of multi-objective bat algorithm to association rule mining. We propose multi-objective binary bat algorithm (MBBA) based on Pareto for association rule mining. This algorithm is independent of minimum support and minimum confidence. To evaluate the association rules mined by MBBA algorithm, we propose a new method to discover interesting association rules without favoring or excluding any measure. Compared with the single-objective BPSO, binary bat algorithm (BBA) and Apriori algorithm, the experimental results on six datasets show that the new algorithm is feasible and highly effective. It can make up the shortage of single objective algorithms and traditional association rule mining algorithms.
Ranking solutions of the population in an evolutionary algorithm that solves a many objective optimization problem is a challenging task which has been vastly studied in recent years. Loss in the hypervolume of the population when a solution is omitted could be a good measure for ranking solutions but calculating this value for high dimensional problems is not tractable. In the selection operator of evolutionary algorithms, the actual hypervolume values are not important, with only a relative knowledge of these values we can select best solutions. In this paper we have proposed a novel method for approximating the ranking induced by the hypervolume indicator. This method is compared to the similar methods and its efficiency and performance is proved via proper experiments. The method is tested using benchmark test problems from the Walking Fish Group problem family which are scalable both in number of variables and objectives.
Krill Herd (KH) optimization algorithm was recently proposed based on herding behavior of krill individuals in the nature for solving optimization problems. In this paper, we develop Standard Krill Herd (SKH) algorithm and propose Fuzzy Krill Herd (FKH) optimization algorithm which is able to dynamically adjust the participation amount of exploration and exploitation by looking the progress of solving the problem in each step. In order to evaluate the proposed FKH algorithm, we utilize some standard benchmark functions and also Inventory Control Problem. Experimental results indicate the superiority of our proposed FKH optimization algorithm in comparison with the standard KH optimization algorithm.
The Particle Swarm Optimization (PSO) is a heuristic optimization technique-based swarm intelligence that can be applied to solving many real-world optimization problems. However, the standard PSO algorithm can easily get trapped in the local optima and has slow convergence speed, and these drawbacks have hindered its further development in all fields. In this paper, a new optimization method based on neighbor heuristic and Gaussian cloud learning is introduced in order to improve the performance of traditional PSO (NHPSO). The NHPSO consists of two main steps. First, by analyzing the relationship among particles in the evolutionary process, a neighbor heuristic mechanism is performed to improve the search efficiency and convergence speed. In addition, a Gaussian cloud learning strategy is introduced to enhance population diversity and balance the global and local search abilities. The performance of the NHPSO is tested using 12 benchmark functions and 6 shifted functions. Results show that NHPSO is superior to the recent variants of PSO in terms of convergence speed, solution accuracy, algorithm efficiency and robustness.
Analyzing different drugs for various purposes is an important issue in the area of computational biology. We categorize the previous computational studies into Individual and Network approaches. While the Individual approach focuses on one specific drug without considering its relationship with other drugs, the Network approach considers also the drugs relationships. In this paper, we apply a Network approach, previously proposed for discovering the relationships among diseases, to drug data. We construct a Human Drug Network (HDN) for 200 different drugs based on functional and structural information available in the PPI network. For evaluating our proposed HDN, first, we analyzed the literature to prove that the proposed HDN is biologically meaningful. Second, we used the HDN to augment the initial prior knowledge of different drugs. As an example of prior knowledge, we considered the initial seed proteins (a set of proteins which are previously known to be drug targets) of each drug. We clustered the HDN nodes using the Markov CLustering Algorithm (MCL) and then, we augmented the seed proteins of each drug based on the cluster it belongs to. In the end, we concluded that our proposed HDN enables us to generate novel hypotheses (in terms of potential drug target proteins) and produce complementary results comparing to existing methods.
Influence maximization in a social network involves identifying an initial subset of nodes with a pre-defined size in order to begin the information diffusion with the objective of maximizing the influenced nodes. In this study, a sign-aware cascade (SC) model is proposed for modeling the effect of both trust and distrust relationships on activation of nodes with positive or negative opinions towards a product in the signed social networks. It is proved that positive influence maximization is NP-hard in the SC model and influence function is neither monotone nor submodular. For solving this NP-hard problem, a particle swarm optimization (PSO) method is presented which applies the random keys representation technique to convert the continuous search space of the PSO to the discrete search space of this problem. To improve the performance of this PSO method against premature convergence, a re-initialization mechanism for portion of particles with poorer fitness values and a heuristic mutation operator for global best particle are proposed. Experiments establish the effectiveness of the SC in modeling the real-world cascades. In addition, PSO method is compared with the well-known algorithms in the literature on two real-world data sets. The evaluation results demonstrate that the proposed method outperforms the compared algorithms significantly in the SC model.