
Review article
Select search scope: search across all journals or within the current journal

The volume of short text data increases rapidly these years. Data examples include tweets and online Q&A pairs. It is essential to organize and summarize these data automatically. Topic model is one of the effective approaches, whose application domains include text mining, personalized recommendation and so on. Conventional models like pLSA and LDA are designed for long text data. However, these models may suffer from the sparsity problem brought by lacking words in short text scenarios. Recent studies such as BTM show that using word co-occurrent pairs is effective to relieve the sparsity problem. However, both BTM and extended models ignore the quantifiable relationship between words. From our perspectives, two more related words should occur in the same topic. Based on this idea, we introduce a model named RIBS, which makes use of RNN to learn relationship. By using the learned relationship, we introduce a model named RIBS-Bigrams, which can display topics with bigrams. Through experiments on two open-source and real-world datasets, RIBS achieves better coherence in topic discovery, and RIBS-Bigrams achieves better readability in topic display. In the document characterization task, the document representation of RIBS can lead better purity and entropy in clustering, higher accuracy in classification.
Due to the rapid growth of web platforms such as blogs, discussion forums, peer-to-peer networks, and various other types of social media, Sentiment Polarity Detection (SPD) (classifying texts by “positive” or “negative” orientation) has become more important and challenging task in recent years. There is a growing need for management and study of SPD not only in English, but also in other languages. The key reason for using Machine Learning (ML) for SPD lies in engineering a representative set of features. This paper explores different (byte, character and word)
Empirical evidence suggests that ensembles with adequate levels of pairwise diversity among a set of accurate member algorithms can significantly outperform any of the individual algorithms. As a result, several diversity measures have been developed for use in optimizing ensembles. We show, however, that there is natural tension between the pairwise diversity of ensemble members and their individual accuracy. While efficient ensembles can be built with stronger forms of diversity, they also suffer in overall accuracy. On the other hand, ensembles built with weaker forms of diversity can be very accurate, but tend to be significantly more computationally expensive. We discuss these findings in light of the notion of diversity space.
A number of graph-parallel computing abstractions have been proposed to address the needs of solving complex and large-scale graph computing. However, unnecessary and excessive communication and state sharing between nodes in these frameworks not only reduce the network efficiency but may also cause decrease in runtime performance. In this paper, we propose a mechanism called LightGraph, which reduces the synchronizing communication overhead for distributed graph-parallel computing abstractions. Besides identifying and eliminating the redundant synchronizing communications in existing systems, in order to minimize the required synchronizing communications LightGraph also proposes an edge direction-aware graph partitioning strategy. This new graph partitioning strategy optimally isolates the outgoing edges from the incoming edges of a vertex. We have conducted extensive experiments using real-world data, and our results verified the effectiveness of LightGraph. For example compared to PowerGraph LightGraph can not only reduce up to 31.5% synchronizing communication overhead for intra-graph synchronizations, but also cut up to 16.3% runtime for PageRank running on Livejournal dataset.
Spatial co-location mining is a useful tool for discovering spatial association patterns of feature sets which are frequently observed together in nearby geographic space. Most of co-location mining techniques aim to find all prevalent co-located feature sets which satisfy a given prevalence threshold. However the result is often large, especially when the prevalence threshold is set low, or long co-location patterns present. Moreover the output has many redundant information which makes it difficult for users to filter useful patterns. This work introduces the problem of mining reduced sets of co-location patterns in order to concisely represent interesting spatial relationship patterns. With aiming two such outputs in the form of maximal and closed co-locations, this paper proposes an algorithmic framework to discover maximal co-location patterns and closed co-location patterns as well as all prevalent co-location patterns, and presents the algorithm details for each pattern discovery. The developed algorithms are correct and complete in finding maximal co-locations and closed co-locations. The experiment result shows that the framework reduces candidate feature sets effectively and finds co-location patterns efficiently.
Time series classification and class imbalance problem are two common issues in a multitude of real-life scenarios. This paper simultaneously explores both issues with deep convolution neural networks (CNNs). Because standard networks treat the majority and minority classes with same class weights, most CNN-based networks fail to classify imbalanced time series. Until recently, there is very little work applying deep learning to imbalanced time series classification (ITSC). Thus, we propose an adaptive cost-sensitive learning strategy to address the ITSC problem. The standard CNN is modified to a cost-sensitive network (CS-CNN), which is able to punish the misclassified samples using a class-dependent cost matrix. Moreover, this cost matrix is automatically updated based on overall class distribution and the CS-CNN’s training performance. The proposed method is extended to FCN, LSTM-FCN and ResNet. It is experimentally tested on five public benchmark UCR datasets and a real-life large volume dataset. Four cost-sensitive CNN-based networks are compared with several data samplers and two traditional ITSC methods. The modified networks are superior in all metrics. Results show that cost-sensitive networks successfully complete the ITSC tasks.
In multi-label classification settings, one of the most common problems is the massive label output space. To alleviate this, some methods opt to exploit label correlations to reduce the output space during prediction. However, these methods sacrifice efficiency or ignore global label correlations. In addition, label imbalances are another problem that is prevalent in multi-label classification. Current methods of correcting for imbalance oftentimes use single-label methods, which fail to consider label correlations. In this paper, we introduce general frameworks that incorporate topic modeling to seamlessly address both problems. We show that these frameworks can allow even the most naïve methods, such as Binary Relevance, to perform similarly to state-of-the-art methods. Furthermore, we show that our frameworks can also adapt state-of-the-art methods to perform better than the methods by themselves.
This study establishes the new results for Cluster Width of probability Density functions (CWD). There are the upper and lower bounds of CWD and the relationships of CWD to other measures in statistical discriminant. The CWD for two and more two probability density functions is determined in the different cases. Based on CWD, we propose a measure called similar coefficient to evaluate the quality of the established clusters. Furthermore, CWD is also used as a criterion to build two algorithms: to determine the suitable number of clusters and to analyse the fuzzy clusters. The numerical examples are given to illustrate the proposed algorithms and to prove their advantages over existing methods.
Reddit is a popular social media website where users can submit content such as direct links and text posts into a forum called subreddit. The average number of new subreddits created reaches 500 per day. Because of the vast and growing number of subreddits, users need to discover and familiarize themselves with all existing communities before submission. In this paper, we propose new feature sets for an online community which are text posts ratio, the average length of text in the post and the domain-specific features. The community recommendation framework is designed and experimented based on Reddit dataset. The framework successfully identifies and collects textual communities by finding their representatives using clustering algorithm namely DBSCAN, then a logistic regression algorithm is applied to recommend a list of communities with high content similarity to a given post. Comprehensive experimental evaluations on Reddit dataset reveal that the proposed framework achieves high precision at 90%.
This research focuses on resource assignment in cooperative energy heterogeneous systems with non-orthogonal multiple access in which cells are powered via a common grid network and alternative energy resources and all base stations have the ability to cover a group of subscribers simultaneously at a specific frequency band. In order to consider the local limitations of alternative energy resources, it was assumed that the alternative energy would be shared among the base stations by the dynamic grid network. In this architecture, resource allocation and user association frameworks should be reconfigured because conventional schemes use orthogonal multiple access. Hence, this paper suggests a novel approach joint optimal power allocation and user association techniques to achieve the maximum degree of energy efficiency for the whole system in which the quality of experience parameters are assumed to be bounded during multi-cell multicast sessions. The solution to the introduced problem in a scenario with fixed transmission power is an improved decentralized algorithm that supplies effective user association framework. The model has been modified to develop joint multi-layered resource control and user association that can distinguish the service pattern in cooperative energy heterogeneous systems with non-orthogonal multiple access to obtain more resource optimality than current approaches. The effectiveness of the suggested approach has been confirmed by the numerical results. Also, the results reveal that non-orthogonal multiple access can provide greater energy efficiency than orthogonal multiple access in heterogeneous wireless networks.
Electricity consumption prediction in smart homes and its effective management are global concerns. One of the most important inventions to assist human living, electricity is used by residential users as well as commercial operations. These users often utilize different electronic devices and sometimes consume fluctuating amounts of electricity, generated from smart-grid infrastructure owned by the government or private investors. However, a repeated imbalance is noticeable between the demand and supply of electricity; these disparities are often brought about by different weather profiles such as temperature, wind speed, dew point, humidity and pressure of the electricity consumption locations. Therefore, effective planning through an intelligent data analysis of the electricity load is needed to enable a sustainable distribution among consumers. Such intelligent analysis and planning are activated by the need to visualize the data and predict future electricity consumption within a short period, considering how weather variables affect predictions. Although a variety of compelling state-of-the-art techniques are used for such predictions, they require data engineering improvement for reducing significant predictive errors in short-term load forecasting (STLF). This research deploys a near-zero cooperative probabilistic scenario analysis and decision tree (PSA-DT) model to address the predictive errors facing state-of-the-art models, and analyses the effect each weather profile has on the cooperative model. The PSA-DT is a machine learning (ML) model based on a probabilistic technique (in view of the uncertain nature of electricity consumption), complemented by a DT to reinforce collaboration between the two techniques. Based on detailed experimental intelligent data analytics (IDA) on residential and commercial data loads, together with multiple weather profiles, the PSA-DT model outperforms state-of-the-art models in terms of accuracy to a near-zero error rate. This implies that its deployment for electricity demand in planning smart homes will be of great benefit to various smart-grid operators and homes.
Short-term traffic flow prediction plays a crucial component in transportation management and deployment. In this paper, a novel regression framework for short-term traffic flow prediction with automatic parameter tuning is proposed, with the SVR being the primary regression model for traffic flow prediction and the Bayesian Optimization being the major method for parameters selection. First, the preprocessing of raw traffic flow is carried out by seasonal difference to eliminate the non-stationary of the data. Then, Support Vector Regression model is trained by the pre-processed data. In order to optimize the model parameters, the generalization performance of SVR is modeled as a sample from a Gaussian process (GP). Bayesian optimization determines the parameters configuration of the regression model by optimizing the acquisition function over the GP. Finally, the optimal short-term traffic flow regression model is constructed through repeated GP update and iteratively multiple training of the model. Experiment results show that the accuracy of proposed method is superior to methods of classical SARIMA, MLP-NN, ERT and Adaboost.