
Editorial
Select search scope: search across all journals or within the current journal

In order to overcome the problems of low accuracy and time-consuming of traditional prediction methods for short-term traffic flow in urban, a prediction methods for short-term traffic flow in urban based on multiple linear regression model is proposed. The corresponding data attributes of short-term traffic flow in urban are selected by traffic operation status, and used as the original data of traffic flow prediction. According to the selected attributes, spatial static attributes data and traffic flow dynamic attributes data are collected, and fault data are identified and repaired. A multiple linear regression model for prediction of short-term traffic flow in urban is constructed to realize the prediction of short-term traffic flow in urban. The experimental results show that, compared with other methods, the average prediction accuracy of the proposed method is as high as 98.48%, and the prediction time is always less than 0.7 s, which is shorter.
In order to overcome the serious errors of wind farm load abnormal fluctuation forecasting results caused by traditional forecasting methods, a wind farm load abnormal fluctuation forecasting method based on probabilistic neural network is proposed in this paper. The probabilistic density is screened out by probabilistic neural network, and the maximum posterior probability density neuron is used as the output to realize wind farm load forecasting. According to the prediction results, a comprehensive severity subordinate function is constructed based on fuzzy reasoning to classify the severity of wind farm anomalies. According to the fuzzy operation rules, the abnormal fluctuation of wind farm load can be warned. The experimental results show that the operation error of the proposed method is only 0.49, the accuracy of early warning is high, and the effective fitting index is up to 0.95, which shows that the proposed method has high practical application value.
In order to improve the accuracy and efficiency of power instability prediction for wind turbines, a power instability prediction method for wind turbines based on fuzzy decision tree is proposed. According to the variation curve of maximum output power, the maximum power of wind turbine is searched and controlled by climbing hill. The maximum power of wind turbine is tracked by the control results. The power fluctuation periodicity rule is obtained based on the fuzzy decision tree. The power instability prediction model of wind turbine is established to realize the power instability prediction. The experimental results show that the proposed method has high effectiveness, the highest prediction accuracy can reach 95.53%, and the maximum prediction time is only 1.8 s, which fully shows that the proposed method is more suitable for the power instability prediction of wind turbines.
Edge computing extends the cloud computing paradigm to the edge of the network to offset the shortcomings of traditional cloud computing in mobile support, low delay, and location awareness. However, traditional complicated encryption algorithms, access control measures, identification protocols, and privacy protection methods are inapplicable to security defense of edge computing given the multisource data fusion characteristics of edge computing, superposition of mobile and the Internet, and resource limitations in storage, computing, and battery capacity of edge terminals. Therefore, a reasonable defense model for complicated dynamic edge computing environment must be established. In accordance with the characteristics of limited resources of edge devices, combined with dynamic game theory, this study proposes an optimal defense strategy model based on a differential game and solves the optimal defense strategy of edge nodes in infinite and finite-horizons. The optimal defense strategy of the edge nodes was simulated under different conditions, and the simulation verifies that the edge nodes can obtain the optimal defense effect with minimum resource consumption when cooperating to form the defense system.
This article has been retracted. A retraction notice can be found at https://doi.org/10.3233/JIFS-219434.
Construction change management is an important part of construction project management. Previous studies usually focus on the various steps of construction change control: analyzing the causes of change, avoiding the risk of change, tracking the process of change, management and feedback of change. This study focused on improving the information flow and organizational relationship of the participants in the process of construction change management, which was reengineered with the use of Building Information Modelling (BIM) technology for better information integration. As BIM technology provides technical means for information transaction and management, a reengineered construction change management process based on BIM technology was proposed to form a more effective way of putting forward, examine, issuing, updating and archiving the information. Both the workflow and the organizational units involved in each work step were reengineered to improve the information flow. To verify the effect of process reengineering, Uncinet, a visual analysis tool for Social network analysis (SNA) was applied to quantitatively analyze and compare the traditional and BIM based organizational structure. The research results demonstrated that BIM was helpful to strengthen organizational coordination and information exchange in construction change management process.
With the increasing global energy crisis, wind power is gradually favored by people for its sustainability and cleanliness. When the wind power is connected with the power system, the problem of grid connection must be overcome. In order to better meet the demand load interactive response in smart grid environment, a multi-layer scheduling two-dimensional operation model was proposed. The mechanism of wind power uncertainty was analyzed. Thus, wind power prediction algorithm was established, and the concept of probability based multi-level scheduling was introduced to improve the adaptability of the model as a whole. In order to verify the reliability and effectiveness of the method, case experiments were carried out. The results show that the method can effectively help wind power consumption when ensuring economic and reliable conditions.
In order to overcome the problem of low fitting between traffic uncertainty prediction results and actual values in existing research methods, a traffic flow uncertainty prediction method based on K-nearest neighbor algorithm is proposed. The original database, classification center database, k-nearest neighbor database and intermediate search database are used to construct the database needed in the prediction process. Based on the database, multivariate linear regression is used to assign weights to state variables, and k-nearest neighbor algorithm and Kalman filter are used to update the weights to adapt to the uncertainties of traffic flow until the predicted values are obtained, and the uncertainties of traffic flow are predicted. The experimental results show that the maximum average absolute error and average relative error of the proposed method are 0.018 and 0.02, respectively. Compared with the traditional method, the proposed method has higher overall prediction accuracy, higher fitting degree, and is feasible.
In order to overcome the problems of low accuracy and long time-consuming in traditional short-term forecasting methods for dynamic traffic flow, a short-term forecasting method for dynamic traffic flow based on stochastic forest algorithm is proposed in this paper. This method chooses short-term forecasting equipment for dynamic traffic flow, eliminates invalid data from the collected data, and normalizes the available data to complete data preprocessing before traffic flow forecasting. A combined forecasting model is established to optimize the output of the pretreatment results and complete the dynamic traffic flow rate forecasting. On this basis, the stochastic forest algorithm is introduced to train the sampling set of flow rate decision tree and generate short-term flow decision tree to realize short-term forecasting of dynamic traffic flow. The experimental results show that the forecasting time of the proposed method is short, always less than 0.5 s, and the forecasting accuracy is high, with more than 97%, so it is feasible.

With the development of the large ship, automated container terminals (ACTs) have serious energy consumption and carbon emission problems, reducing the loading and unloading time of ships can ease energy consumption, improve the working efficiency and service level of automated terminals. This paper studies the integrated scheduling problem of the gantry cranes (QCs), automated guided vehicles (AGVs) and automated rail-mounted gantry (ARMG) in automated terminal. According to the loading and unloading operation mode, we build the mixed integer programming model with the goal of minimizing the ship loading and unloading time, and through various algorithms of heuristic and hybrid improved to solve this problem, it proves the effectiveness of the model to obtain optimized scheduling scheme by numerical experiments, and comparing the different performance of algorithms, the results show that the hybrid GA-PSO algorithm with adaptive auto tuning is superior to other algorithms in terms of solution time and quality, which can effectively solve the problem of integrated scheduling to save the energy of automated container terminal.
In order to effectively improve the accuracy of related analysis models in the application of government risk investment, a government risk investment prediction model based on fuzzy clustering discrete algorithm is put forward in this paper. First of all, government risk investment problem is analyzed. Based on Markowitz theory, the general government risk investment model is considered, and the market value constraint and the upper bound constraint are combined to improve the government risk investment model and obtain the mixed constraint government risk investment model. Secondly, the fuzzy clustering discrete algorithm is introduced in the analysis process of government venture investment model, and it is used to solve the mixed constraint analysis model of government venture investment. In addition, to further improve the performance of discrete algorithm based on fuzzy clustering in the model solving process, automatic contraction and expansion of factors is used to carry out adaptive learning of related parameters based fuzzy clustering discrete algorithm, and improve the convergence of the algorithm. Finally, the simulation experiments on some stock samples of investment sector show that the algorithm in this paper can obtain more ideal government venture investment schemes, so as to reduce investment risk and obtain greater investment returns.
Vehicular ad hoc networks play an important role in current intelligent transportation networks, which have attracted much attention from academia and industry. Vehicular networks can be implemented by Long-Term Evolution Advanced (LTE-A) networks, which have been formally defined in a series of standards by third-generation partnership projects (3GPP). Abundant challenges exist in the authentication processes in LTE-A-based vehicular networks. This paper aimed to improve the security functionality of these vehicular networks by proposing a secure and efficient group authentication and privacy-preserving scheme for vehicular networks based on fuzzy system: the group authentication and privacy-preserving level (GAPL). Compared with existing schemes, the proposed scheme can greatly reduce the number of control message transmissions from mass vehicular equipment (VEs) to the network and substantially avoid overhead in LTE-A-based vehicular networks. Privacy-preserving levels are established to protect VE privacy in authentication. Furthermore, the scheme contains security functions, including privacy preservation, non-frameability and non-repudiation verification.
Influenced by national policies and macro-economic environment, large domestic enterprises is actively promoting strategic transformation to enhance their core competitiveness, and performance evaluation of enterprises’ innovation capacity has become a hot topic in recent years. This paper proposes a performance evaluation method of enterprises’ innovation capacity based on deep learning fuzzy system model and convolutional neural network analysis of innovation network. First of all, on account of the characteristics of breakthrough innovation and drawing on the traditional innovation performance evaluation model, this paper constructs a breakthrough innovation performance evaluation index system for enterprises from the six dimensions of main resource input, technology out-turn, process management, product performance, social value and commercial Value. Secondly, the introduction of machine learning of fuzzy convolutional neural network to assess the advancement execution of enterprises is of great significance for enterprise managers to find out the problems and causes of enterprises’ innovation, optimize the allocation of enterprises’ resources and further improve the innovation performance of enterprises. The experimental results show to verify the adequacy of the algorithm.
A control strategy of permanent magnet-oriented field synchronous motor based on intelligent fuzzy control system and generalized predictive control with non-linear identification is proposed to develop the effectiveness of the controlling method of constant magnet-oriented field synchronous motor, the accessor can be split into stabilization control part and intelligent control part. The input of traditional feedback control is used as the stabilization control part, while the feed-forward is incorporated into the intelligent part to compensate for the uncertainties of repetitive load torque and model parameters. The proposed feed forward compensation term uses simple learning rules without any load torque disturbance observer. The additional learning feed forward term does not require information about motor parameters and load torque values, it is insensitive to load torque uncertainty and model parameters, and does not need to identify the system model. With that, the solidness and intermingling confirmation of the proposed control framework reaction is given. The exploratory outcomes demonstrate that the proposed technique has littler speed overshoot list, and the heap torque against aggravation capacity list is improved by over 30%.

In order to study the water flow movement of the Yazidang Reservoir, this paper generates the initial terrain for the researched water area with the image stitching technology and image edge detection technology, establishes a 3D
Fake online reviews are so prevalent that e-commerce platforms attempt to control it from affecting the trustworthiness between buyers and sellers. The issue has also attracted sporadic scholarly endeavor to understand this new field. To address this issue, we propose a new model to examine three interrelated stakeholders of e-Commerce platforms: experienced buyers, future buyers and the online sellers in terms of purchasing behaviors and sales with three objectives. Experienced buyers influence future consumers’ behaviors and increase sales from sellers. Using data collected from the largest online e-commerce platform in China, we test relevant hypotheses. Our findings show that experienced buyers and their positive reviews increase future buyers’ purchasing and promote corporate sales. These findings contribute knowledge to the online feedback mechanism and literature on fake review studies. This study also provides a novel method to help buyers avoid fake online review from a market structure perspective.
Over the years protein interaction and prediction of membrane protein have been a pivotal research area for all researchers. For both prokaryotes and eukaryotes Adenosine Triphosphate-(ATP) binding cassette (ABC) genes plays a significant role. In our analysis, we concentrate on human part of ABC genes. In case of living organisms transport of precise molecules across lipid membranes has been treated as vital part and for that reason a bigger transporter is required to carry out the molecules. Here ABC transporter families are evolved to transport the specific molecules such as sugars, amino acid, peptides, proteins, ions etc. within the plasma membrane. As we know another important component of human being is cholesterol, which is a major component in cell membrane and its main functions are to maintain integrity and mechanical stability. Each and every time, membrane cholesterolsareinteracted with membrane protein in both N-C terminuses and target valid sequence(s) which has relevance in human diseases. In this manuscript we have applied Fuzzy C-Means (FCM) with Support Vector Machine (SVM) algorithm for prediction of cellular cholesterol with ABC genes. Our experiments have been performed well using ABCdata set.
For the unsupervised learning based clustering algorithm, the intrusion detection rate is low, and the training sample based on supervised learning clustering algorithm is insufficient. A semi-supervised kernel fuzzy C-means clustering algorithm based on artificial fish swarm optimization (AFSA-KFCM) is proposed. Firstly, the kernel function is used to change the distance function in the traditional semi-supervised fuzzy C-means clustering algorithm to define a new objective function, thus improving the probabilistic constraints of the fuzzy C-means algorithm. Then, the artificial fish swarm algorithm with strong global optimization ability is used to improve the KFCM sensitivity to the initial cluster center and easy to fall into the local extremum, thus improving the convergence speed and improving the classification effect. The test results in the Wine and IRIS public datasets show that the AFSA-KFCM clustering algorithm is superior to the traditional algorithm in clustering accuracy and time efficiency. At the same time, the experimental results in KDDCUP99 experimental data show that the algorithm can obtain the ideal detection rate and false detection rate in intrusion detection.
In order to overcome the problems of long encrypting time, low information availability, low information integrity and low encrypting efficiency when using the current method to encrypt the communication information in the network without constructing the sequence of communication information. This paper proposes a network communication information encryption algorithm based on binary logistic regression, analyses the development of computer architecture, builds a network communication model, layers the main body of information exchange, and realizes the information synchronization of device objects at all levels. Based on the binary Logistic regression model, network communication information sequence is generated, and the fusion tree is constructed by network communication information sequence. The network communication information is encrypted through system initialization stage, data preparation stage, data fusion stage and data validation stage. The experimental results show that the information availability of the proposed algorithm is high, and the maximum usability can reach 97.7%. The encryption efficiency is high, and the shortest encryption time is only 1.9 s, which fully shows that the proposed algorithm has high encryption performance.
In order to overcome the problems of poor accuracy and high complexity of current classification algorithm for non-equilibrium data set, this paper proposes a decision tree classification algorithm for non-equilibrium data set based on random forest. Wavelet packet decomposition is used to denoise non-equilibrium data, and SNM algorithm and RFID are combined to remove redundant data from data sets. Based on the results of data processing, the non-equilibrium data sets are classified by random forest method. According to Bootstrap resampling method with certain constraints, the majority and minority samples of each sample subset are sampled, CART is used to train the data set, and a decision tree is constructed. Obtain the final classification results by voting on the CART decision tree classification. Experimental results show that the proposed algorithm has the characteristics of high classification accuracy and low complexity, and it is a feasible classification algorithm for non-equilibrium data set.
In order to overcome the problems of invulnerability and low communication efficiency when analyzing network communication instability with current methods, this paper proposes a modeling method of network communication instability based on K-means algorithm. The network element nodes are generated by clustering idea, and the initial communication topology is constructed. K-means algorithm is used to optimize the initial communication model, build a comprehensive mathematical model of network communication, and solve the model to realize the optimization of communication model. The network efficiency function is used to further quantify the network invulnerability, and the function is used to find the most vulnerable nodes in the network, and strengthen them to achieve efficient control of network invulnerability. The experimental results show that the model has strong invulnerability, up to 99.9%, high communication efficiency and coverage, and the maximum communication delay is only 0.35 s. It is a feasible network communication model.
In order to overcome the inaccuracy of current research results of traffic flow prediction, this paper proposes a prediction method for traffic flow with small time granularity at intersection based on probability network. This method takes one minute as time granularity, collects traffic data such as cross-section flow, section traffic flow velocity data, traffic density, road occupancy, section delay and steering ratio by using RFID technology, and analyzes and processes the data. By introducing Bayesian network in probabilistic network and combining K-nearest neighbor method, historical data and predicted traffic flow state are classified to realize the prediction of traffic flow with small time granularity at intersections. The experimental results show that this method has high prediction accuracy and reliability, and is a feasible traffic flow prediction method.
Social media is becoming more and more closely related to the real life. More and more netizens choose to obtain news and publish notice through social networks. Such huge amount of social media information generated by these users contains a lot of information related to hot topics and events. At the same time, problem of information overload has posed a challenge for people to use the information. It has become an important research issue to discover and track hot events and topics automatically from mass social media data. On the one hand, the short, highly noisy and real-time features of the social media data bring challenges to the discovery and tracking methods of traditional hot issues. On the other hand, the social media data contains abundant information of geography, time, and social relations, which brings great convenience to relevant researches. Based on these features of the social media data, this paper makes a deep study on the discovery, extraction, and tracking of hot issues in the social media based on fuzzy system theory and the word vector semantic clustering.
On the basis of FHWA model of the Federal Highway Administration and the combination with the geographic information system (GIS) and Fuzzy intelligent control system, the group independently researches and develops a simulation and evaluation system for the traffic noise in the urban road. This system is able to simulate the influence of traffic source, point source, and arbitrary shape area source on the urban sound field environment. It is combined with the noise radiation and the communication model, and the occlusion and attenuation by the buildings and forest belts on the traffic noise have been considered. It can calculate the traffic noise in urban areas and directly render the predicted results on the GIS map, and form a traffic noise map, which visually and clearly displays the pollution degree and distribution map of the traffic noise in urban areas. The noise maps of Guangzhou inner ring roads and Zhujiang New Town are drawn to provide scientific decision-making basis for the control of urban traffic noise pollution.
With the increasing amount of information on the Internet, data storage management tends to be distributed. In distributed storage environment, users pay more and more attention to the timeliness of user interaction experience and the reliability of information interaction. However, the efficiency of users is often limited by the efficiency of data communication between distributed sites. One of the important goals of distributed data management is to improve the efficiency of data transmission and ensure the reliability of data transmission. Block chain technology is one of the emerging technologies supporting the development of management information system; it provides a solution for the storage, verification, transmission and communication of the distributed data. This paper focuses on solving the problem of block chain data transmission, and studies it from three aspects: improving the efficiency of data communication, ensuring the reliability of transmission, and improving the fairness of service, and different block chain data communication performance optimization strategies are proposed under the constraints of node communication capability, node trust, weight, priority of service request and other influencing factors.



It is of great significance to explore the prospects of English intelligent learning in the field of basic education and to understand the current status and practical needs of mobile learning technology. Based on the intelligent English learning and teaching needs, this study constructs an intelligent English learning system based on improved neural network. The system uses wavelets to replace the neurons in the traditional neural network, establishes the connection between the wavelet transform and the network coefficients through affine transformation, and uses the wavelet neural network to correct the network parameters, so as to avoid the problem of being partially optimal due to the sensitivity of the initial value to some extent. Moreover, the English learning system adopts a mature C/S structure, and the system is divided into two parts and four layers. In addition, this study designed the experiment to analyze the performance of the English learning system, set the experimental group and the control group to conduct practical analysis. The research results show that the system constructed in this paper has certain effects and can provide theoretical reference for subsequent related research.
The current music education information has certain information loss in the real-time transmission process, which leads to poor educational effect. The research content of this thesis is based on FPGA video image acquisition and processing system. At the same time, this research mainly uses FPGA as the platform to realize and simulate video acquisition, transformation, storage, display and transmission. This research solves the problem of long-distance real-time transmission of high-definition video stream, and in order to improve the subjective quality of the recovered video at the receiving end, the algorithm processing is added to reduce the blockiness phenomenon in the video displayed at the receiving end. Finally, the validation test was designed to validate the research perspective. The experimental results show that the system works stably and realizes the functions of video image acquisition, conversion, display and transmission, and achieves the design goal. However, due to time constraints and other conditions, the entire design needs further improvement, so that it can provide a basis for the development of subsequent music education information communication technology.
Online education has become an important way of learning English at present, and English vocabulary teaching can improve the efficiency of English vocabulary teaching through target visual detection. However, from the existing research, it can be seen that there are still some shortcomings in English vocabulary recognition. In order to improve the English vocabulary recognition effect, based on machine learning recognition technology, this study combines English vocabulary recognition needs of online education to construct an English vocabulary detection model based on convolutional neural network. The model takes the word’s overall feature as the feature extraction principle and adopts the analysis and extraction of the joint segment feature. Moreover, it discards the complicated process of first dividing a single letter and then performing feature extraction and recognition. In addition, this study design example tests to perform algorithm performance analysis. The experimental results show that the proposed algorithm model has certain effects, and it can be used as an auxiliary algorithm for online English vocabulary teaching.
When the English teaching text is regarded as the ontology, it must involve how to describe the attribute effectively. However, in the current research, the research on the automatic extraction of labels for English teaching texts is still insufficient. Intelligent English teaching has become an inevitable trend in the development of future English teaching models, so it is necessary to cooperate with intelligent text recognition technology. Based on SVM, this study applies convolutional neural network algorithm to text recognition of English teaching content, and effectively recognizes text features. After feature extraction, the original text content has been changed into data that the machine can directly identify and analyze, and semantic analysis is performed. In order to verify the performance of the algorithm, the performance of the algorithm was analyzed by example verification. It can be seen from the results that the proposed method has a certain accuracy rate and can be applied to the text recognition classification of English teaching content and can provide reference direction for related research.
At present, the automatic attendance mode of distance education is not conducive to the confirmation and analysis of information after class. In order to study the effective automatic recognition algorithm of remote education classroom, this study takes the educational classroom of intelligent innovation and entrepreneurship of Internet + as an example for analysis. Moreover, this paper adopts facial features as the basis of recognition, establishes corresponding positioning points, and constructs precise positioning methods for real-time feature capture. At the same time, the ASM algorithm is used to extract facial features, and the algorithm is improved to improve the extraction effect. In addition, this paper proposes Gabor-wavelet packet set and Gabor beamlet set for auxiliary recognition, which improves the recognition rate. Finally, this paper designs experiments to analyze the performance of the algorithm of this study. The results show that the proposed algorithm has certain practical effects and can provide theoretical reference for subsequent related research.
In multimedia English teaching, learners face such an indifferent computer screen without emotion and feel the fun of interaction and emotional stimulation, which will cause resentment and affect the learner’s learning effect. In order to improve the efficiency of multimedia English teaching, aiming at the lack of emotion in multimedia English education, this study proposes an intelligent network teaching system model based on deep learning speech enhancement and facial expression recognition. Moreover, this study uses emotional calculation as the theoretical basis and uses facial expression recognition as the core technology to judge and understand the emotional state by capturing and recognizing the facial expressions of online learners. In addition, this study has carried out experimental tests on the effect of the identification method of this paper and verified that the method has good detection effect on the real smile micro-expressions through two sets of experiments and can provide theoretical reference for subsequent related research.
The current online education platform has gradually replaced the traditional teaching mode and has become an efficient teaching method. Flipping classroom is a new teaching mode under the background of the rapid development of information technology. It is also an important way of multimedia network teaching. However, compared with the traditional teaching mode, teachers in the online teaching platform cannot judge the students’ psychological activities through the students’ state of mind, and they can grasp the students’ learning status through the teaching process. Based on this, based on the cloud computing platform, this study improves the data transmission effect and improves the algorithm according to the learning process of the online education platform. Moreover, this study combines support vector machine to construct a student state recognition system suitable for online education platform and conducts algorithm performance analysis through experiments. In addition, this study uses MKmeans algorithm, Kmeans algorithm and improved Kmeans algorithm, that is, K-mediods and Xmeans algorithm to compare the pre-processed final data sets. The research results show that the proposed algorithm is suitable for network teaching platform and has certain practical effects.
At present, applying image recognition technology to promote English teaching is a kind of teaching innovation that meets the needs of the times. Therefore, based on machine learning neural network and image super-resolution, this study conducted an innovative analysis of English teaching mode. This paper combined the current situation of English teaching classroom to study and analyze English classroom, combined classroom characteristics as the basis of English teaching innovation and constructed a feature recognition model suitable for current English teaching status. Moreover, this paper formed an initial high-resolution image for low-resolution image reconstruction by sparse representation method, and then established a mixed sample spine regression model to re-estimate the high-frequency components of the initial high-resolution image to realize various behavioral characteristics of students in English teaching classroom. In addition, this article builds a verification test. The research shows that the proposed algorithm has certain effects and can provide theoretical reference for subsequent related research.
Performance appraisal in business administration has a great impact on social and economic development, so a sound performance appraisal system should be established. Moreover, in the information age, scientific methods are needed to improve business management performance. Based on this, this study links artificial intelligence with convolutional neural networks, and builds a corresponding performance research model based on actual conditions. When building the model, this paper selects the data width of 8Bit and 32 data per line, and shifts storage 2 rows, and sets the read/write enable signal to be half of the clock signal. In addition, the image matrix of the input image subjected to nonlinear processing by the excitation function ReLU will exhibit sparsity. Finally, combined with the model and data constructed in this study, the model is validated and the relevant strategies for performance evaluation are obtained.
It is of great theoretical significance and practical value to analyze the characteristics of users and behaviors in social networks, to study the personalized recommendation algorithms of users, to explore the inherent laws of event development, and to predict the movement of information or opinions. This paper analyzes the Weibo behavior through machine learning and cloud computing technology. Moreover, this paper studies and analyzes traditional network algorithms, and proposes a microblog recommendation algorithm based on statistical features. At the same time, the research content of this paper focuses on microblog contents, user characteristics, user preferences, and influence levels. The algorithm has simple structure and strong computing performance and performs feature data mining through cloud computing big data method, which is suitable for online mining microblog behavior. In addition, the performance of the algorithm was analyzed by design comparison experiments. The research indicates that the research algorithm proposed in this paper has certain advantages, which can be applied to network behavior analysis mining, and can provide theoretical reference for subsequent related research.
The transfer of scientific and technological achievements is an inevitable stage in the application of science and technology to the process of productivity. This process is accompanied by various influencing factors. How to eliminate the influence of adverse influence factors on the transformation of technology into productivity is crucial to the development of social productive forces. Based on this, from the perspective of deep learning, this study builds a technology transfer transformation platform through deep learning combined with data mining technology and analyzes the method in detail. On this basis, this paper takes a city as an example to analyze the platform of scientific and technological achievements transfer. In addition, by collecting existing data as system input and data mining analysis, this paper summarizes the advantages, disadvantages, opportunities and threats of the city’s enterprises in the transformation of results and proposes corresponding countermeasures. The example verification shows that the method proposed in this study has certain practical effects and can provide theoretical reference for subsequent related research.
At present, English teaching does not play the role of a smart classroom, and it is difficult to grasp the student status and characteristics in real time in actual teaching. Based on this, starting from the video image and static image and the actual situation of English classroom teaching, this study, based on the convolutional neural network and random forest algorithm, performs static image human behavior recognition under different image representation conditions, and studies the influence of background information of image and spatial distribution information of image features on recognition accuracy. Then, based on the similarity between different behavior classes, a static image human body behavior recognition method based on improved random forest is proposed. In addition, through theoretical research, an algorithm model that can identify the characteristics of English classrooms is constructed, and the static and dynamic images of English teaching are taken as an example to conduct experimental analysis. The research shows that the proposed method has certain effects and can provide theoretical reference for subsequent related research.
The image content retrieval can effectively promote the development of the entire industry. At present, sports competition is becoming more and more fierce, and the requirements for image content retrieval are getting higher and higher. In this paper, research has been carried out on image descriptor generation, image feature quantization and coding, accurate nearest neighbor cluster center fast search, multi-dimensional inverted index construction and fast retrieval. Moreover, based on deep learning, this paper constructed an effective detection algorithm for the characteristics of sports images, and compared the image shape and color as examples. It can be seen from the comparative study that the research method of this paper can effectively reduce the size of the candidate set of query results without affecting the accuracy of the query, which is of great significance for improving the speed of image query and has certain significance for promoting the development of sports public industry.
With the continuous development of science and technology, computer-aided teaching has become a common mode of school teaching. From the current situation, it can be seen that the current computer-aided teaching mostly replaces the traditional teaching mode with multimedia, and does not play the role of functional teaching, and teachers cannot effectively grasp the students’ psychological thoughts in teaching. Based on this, this study combines machine learning prediction and artificial intelligence KNN algorithm to actual teaching. Moreover, this study collects video and instructional images for student feature behavior recognition, and distinguishes individual features from group feature recognition, and can detect student expression recognition in detail. In addition, this study designed a case study to analyze the performance of the algorithm. From the experimental results, it can be seen that the proposed algorithm has certain effects and can be used as an algorithm to assist the teaching process and can provide theoretical reference for subsequent related research.
If there are more external interference factors in the process of intelligent recognition in English, the recognition accuracy will be greatly reduced. It is of great academic value and application significance to deeply study feature recognition of English part-of-speech and realize automatic image processing of English recognition. Based on unsupervised machine learning and image recognition technology, this study combines the actual factors of English recognition to set the corresponding influencing factors and proposes a reliable method to identify multi-body rotating characters. This method utilizes the principle of the periodic characteristics of the trajectory rotation on the feature space. Moreover, this study conducts a comparative analysis of recognition accuracy by comparative experiments. In addition, this paper analyzes the recognition principles of 4 fonts in detail. The research results show that the proposed method has certain effects and can provide theoretical reference for subsequent related research.
Task degree has become one of the important indicators to measure students’ English learning intensity and learning quality, and the difference in task degree has different effects on students’ English learning. In order to realize the task recognition of English classroom teaching, combined with the characteristics of deep learning, this study combines the actual situation of English classroom teaching to analyze, and distinguishes characters through student positioning and feature recognition. Moreover, this paper combines the characteristics of English learning scoring to judge students’ learning situation, and designs a shallow convolutional neural network based on TensorFlow architecture for identifying images and uses GPU training acceleration to solve the problem of training time-consuming in the face of large data volume. In addition, the task results feedback is evaluated by scoring method, and the performance of the algorithm is analyzed by experiments. By setting the category of sensitive targets, this paper can perceive the results according to the target location and mark the sensitive targets in the input scene image. The research results show that the method proposed in this paper has certain effects.
The rise of the cloud computing model has resulted in more than terabytes of data being stored in the cloud platform every day on the Internet. Mining valuable information from these massive data has become an emerging industry direction, but the current Intrusion-detection system (IDS) has been unable to adapt to large-scale log information mining. Therefore, an association rule mining algorithm based on MapReduce parallel computing framework is proposed. Firstly, the frequent itemsets mining algorithm Apriori is analyzed, and the MapReduce model is used to parallelize and improve it to more efficiently complete the mining of frequent itemsets. Secondly, the parallel Apriori is designed to run on IDS. Finally, the simulation experiment was carried out by building an open source cloud computing framework Hadoop cluster. Finally, the simulation experiment was carried out by building an open source cloud computing framework Hadoop cluster. The results show that the proposed method has higher detection efficiency when processing massive data, and requires less processing time.
Speech Emotion Recognition (SER) has been widely used in many fields, such as smart home assistants commonly found in the market. Smart home assistants that could detect the user’s emotion would improve the communication between a user and the assistant enabling the assistant to offer more productive feedback. Thus, the aim of this work is to analyze emotional states in speech and propose a suitable algorithm considering performance verses complexity for deployment in smart home devices. The four emotional speech sets were selected from the Berlin Emotional Database (EMO-DB) as experimental data, 26 MFCC features were extracted from each type of emotional speech to identify the emotions of happiness, anger, sadness and neutrality. Then, speaker-independent experiments for our Speech emotion Recognition (SER) were conducted by using the Back Propagation Neural Network (BPNN), Extreme Learning Machine (ELM), Probabilistic Neural Network (PNN) and Support Vector Machine (SVM). Synthesizing the recognition accuracy and processing time, this work shows that the performance of SVM was the best among the four methods as a good candidate to be deployed for SER in smart home devices. SVM achieved an overall accuracy of 92.4% while offering low computational requirements when training and testing. We conclude that the MFCC features and the SVM classification models used in speaker-independent experiments are highly effective in the automatic prediction of emotion.
With the promotion of opinion leader’s impact on online purchase intention, the problem of how to measure the characteristics of opinion leader, the characteristics of opinion leader’s recommendation information and the influence of consumers’ characteristics on purchase intention is becoming more and more urgent. Based on numbers of popular scales, this paper designs the questionnaire items for the variables of professional knowledge, product involvement, visual cues, interactivity, functional value and trust involved in the opinion leader influence model, and forms the initial scale. On this basis, with the help of small-scale interviews, small sample pre-test and large sample test, trust and purchase intention fail to pass the validity test. Through correlation coefficient analysis, some questions with lower coefficient value are eliminated, and then the final scale with good reliability and validity is obtained.
In order to overcome the problems of long execution time and low parallelism of existing parallel random forest algorithms, an optimization method for parallel random forest algorithm based on distance weights is proposed. The concept of distance weights is introduced to optimize the algorithm. Firstly, the training sample data are extracted from the original data set by random selection. Based on the extracted results, a single decision tree is constructed. The single decision tree is grouped together according to different grouping methods to form a random forest. The distance weights of the training sample data set are calculated, and then the weighted optimization of the random forest model is realized. The experimental results show that the execution time of the parallel random forest algorithm after optimization is 110 000 ms less than that before optimization, and the operation efficiency of the algorithm is greatly improved, which effectively solves the problems existing in the traditional random forest algorithm.
At present, the teaching of architectural art in China is still relatively traditional, and there are still some problems in the actual teaching. Based on this, this study combines the Naive Bayesian classification algorithm with the fuzzy model to construct a new architectural art teaching model. In teaching, the Naive Bayesian classification algorithm generates only a small number of features for each item in the training set, and it only uses the probability calculated in the mathematical operation to train and classify the item. Moreover, by combining the fuzzy model, the materials needed for architectural art teaching can be quickly generated, and the teaching principles and implementation strategies of architectural art are summarized. In addition, this paper proposes an attribute weighted classification algorithm combining differential evolution algorithm with Naive Bayes. The algorithm assigns weights to each attribute based on the Naive Bayesian classification algorithm and uses differential evolution algorithm to optimize the weights. The research shows that the method proposed in this paper has certain effect on the optimization of architectural art teaching mode.
Temporal information is crucial in knowledge extraction. Being able to locate events in a timeline is necessary to understand the narrative behind every text. To this aim, several temporal taggers have been proposed in literature –nevertheless, not all languages received the same attention. Most taggers work only for English texts, and not many have been developed for other languages. Also the scarcity of annotated corpora in other languages notably hinders the task. In this paper we present a new rule-based tagger called
In this work, we report the results of our experiments on the task of distinguishing the semantics of verb-noun collocations in a Spanish corpus. This semantics was represented by four lexical functions of the Meaning-Text Theory. Each lexical function specifies a certain universal semantic concept found in any natural language. Knowledge of collocation and its semantic content is important for natural language processing, as collocation comprises the restrictions on how words can be used together. We experimented with word2vec embeddings and six supervised machine learning methods most commonly used in a wide range of natural language processing tasks. Our objective was to study the ability of word2vec embeddings to represent the context of collocations in a way that could discriminate among lexical functions. A difference from previous work with word embeddings is that we trained word2vec on a lemmatized corpus after stopwords elimination, supposing that such vectors would capture a more accurate semantic characterization. The experiments were performed on a collection of 1,131 Excelsior newspaper issues. As the experimental results showed, word2vec representation of collocations outperformed the classical bag-of-words context representation implemented in a vector space model and fed into the same supervised learning methods.
A drug name could be confused because it looks or sounds like another. Nevertheless, it is not possible to know a priori the causes of the confusion. Nowadays, sophisticated similarity measures have been proposed focused on improving the score of the detection. However, when a new drug name is proposed, the Federal Drug Administration (FDA) only can reject or accept the drug name based on this value. This paper not only improves the detection of confused drug names by integrating the strengths of different similarity measures but also the orthographic and phonetic knowledge of these measures are used to give an a priori explanation of the causes of confusion. In this paper, a novel measure that integrates 24 individual measures is developed for this problem. With our proposal, each individual measure contributes to this problem. Finally, we present examples of how our proposal is used for explaining the causes of the confusion which could assist to the FDA to accept or reject a new drug name or to know the confusion causes of previously reported cases.
This work presents an experimental study on the task of Named Entity Recognition (NER) for a narrow domain in Spanish language. This study considers two approaches commonly used in this kind of problem, namely, a Conditional Random Fields (CRF) model and Recurrent Neural Network (RNN). For the latter, we employed a bidirectional Long Short-Term Memory with ELMO’s pre-trained word embeddings for Spanish. The comparison between the probabilistic model and the deep learning model was carried out in two collections, the Spanish dataset from CoNLL-2002 considering four classes under the IOB tagging schema, and a Mexican Spanish news dataset with seventeen classes under IOBES schema. The paper presents an analysis about the scalability, robustness, and common errors of both models. This analysis indicates in general that the BiLSTM-ELMo model is more suitable than the CRF model when there is “enough” training data, and also that it is more scalable, as its performance was not significantly affected in the incremental experiments (by adding one class at a time). On the other hand, results indicate that the CRF model is more adequate for scenarios having small training datasets and many classes.
Since a drug name goes through different communication means and circumstances when it is prescribed, written, advertised, listened to, searched and administered; it tends to be confused with similar drug names that Look-Alike and Sound-Alike (LASA). LASA drug names have caused costs and damage to health. For this problem, the institutions of the United Kingdom, Canada, and the United States have implemented programs for several decades to report lists of confusing drug names pairs. Thanks to these kinds of list, it has been possible to propose new models to identify confusing drug names in English and are used to reject new drug name proposals or to alert when a confusing drug name is being dispensed. However, countries such as Spain also have published a list with the Spanish LASA drug names, and it is not clear enough whether the models previously proposed for the drug names in English are useful for the list in Spanish or if it is necessary to adjust and update them for the Spanish language. This paper focuses on updating and improving the identification of LASA drug names in Spanish. First, we update the state-of-the-art by evaluating all the individual similarity measures proposed previously and all the models that combine these measures with the list in Spanish. Second, we updated the models with new individual measures and then adjusted them with the list in Spanish to improve the identification of LASA drug names in Spanish. After that, 25 individual similarity measures and 8 models to identify confused drug names in Spanish are compared to obtain the best result and conclusions.
In this paper, we propose the
In this research work, we propose a rule based approach for the automatic extraction of UML diagram from the unstructured format of software functional requirements. The existing work provides decent results for active sentences and positive sentences but the challenge in our work is to automatic extract class diagram elements from passive voice type sentences and negative sentences. Furthermore, there is scope to do more research in extraction process using multi-word terms. Thus, we have endeavored to automatic extract the class diagram elements by overcoming these challenges. The methodology uses the Stanford CoreNLP Tools along with Java for the practical implementation of formulated rules. Our approach has proved that without supplant the human being and their decision making, one could reduce the human effort while designing functional requirements. Several case studies were performed to compare class diagrams generated by our methodology to the ones created by experts. Our methodology outperforms the existing work and provides impressive Average completeness (0.82), Average correctness (0.92) and Average redundancy (0.15). Results show that class diagram elements extracted by our methodology are precise as well as accurate and hence, in practice, such class diagrams would be a good preliminary diagram to converge towards to precise and comprehensive class diagrams.
Automatic validation of compositionality vs non-compositionality is a very challenging problem in NLP. A very small number of papers in literature report results in this particular problem. Recently, some new approaches have arised with respect to this particular linguistic task. One of these approaches that have called our attention is based on what authors call “lexical domain”. In this paper, we analyze the use of Pointwise Mutual Information for constructing thesauri on the fly, which can be further employed instead of dictionaries for determining whether or not a given phraseological unit is compositional or not. The experimental results carried out in this paper show that this dissimilarity measure (PMI), can effectively be used when determining compositionality of a given verbal phraseological unit. Moreover, we show that the use of thesauri improves the results obtained in comparison with those experiments employing dictionaries, highlighting the use of self-constructed lexical resources which are, in fact, taking advantage of the same vocabulary of the target dataset.
Translation has been one of the oldest problems in natural language processing. Despite its age, it is still one where there is a tremendous scope for improvement and creativity; the quantity and quality of research in it is testament to that fact. The subfield of primarily using deep neural networks for translation has recently started to gain traction. Many techniques have been developed using deep encoder-decoder networks for bilingual translation using both parallel as well as non-parallel corpora. There is a lot of potential in applying concepts such as bilingual embeddings to create generic translation architecture, which doesn’t need huge parallel corpora to train. These ideas are particularly pertinent in the case of Indic languages, where it is generally difficult to obtain such corpus. In this paper, we try to adapt some of newest techniques in autoencoder networks and bilingual embeddings to the task of translating between English and Hindi. The models considerably outperform state of the art translating systems for these languages.
The argumentation in academic writings is necessary to clearly communicate the ideas of the students. The relations between argumentative components are an essential part since this shows the contrast or support of the presented ideas. In this paper, we present two approaches to relation identification between pairs of components. In the first, we detect initially which components are related, to later classify them in support or attack relation. In the second approach, we identify directly which components have a support relation. For these approaches, we employed machine learning techniques with representations of several lexical, syntactic, semantic, structural and indicator features. Experiments in argumentative sections of academic theses showed that the models achieve encouraging results solving the task, and revealing the argumentative structures prevailing in student writings.
Despite advances in medical safety, errors related to adverse drug reactions are still very common. The most common reason for a patient to develop an adverse reaction to a medication is confusion over the prescribed medication. The similarity of drug names (by their spelling or phonetic similarity) is recognized as the most critical factor causing medication confusion. Several studies have studied techniques for the identification of confusing medications pairs, the most important of which employ techniques based on similarity measures that indicate the degree of similarity that exists between two drugs names. Although it generates good results in the identification of confusing drug names, each of the similarity measures used detects to a greater or lesser degree of similarity that exists between a pair. Recent studies indicate that the optimized combination of several similarity measures can generate better results than the individual application of each one. This paper presents an optimized method of combining various similarity measures based on symbolic regression. The obtained results show an improvement in the identification of confusing drug names.
In this work, we present a model for the automatic generation of written dialogues, through the use of grammatical inference. This model allows the automatic recognition of grammars from a set dialogues employed as a training set. The inferred grammars are then used to generate templates of responses within the dialogues. The final objective is to apply this model in a specific domain dialogue system that answers questions in Spanish with the use of a knowledge base. The experiments carried out have been performend using the DIHANA project corpus which contains dialogues written in Spanish about schedules and prices of a rail system.
When people communicate, we often face situations where decisions have to be made, regardless of silence of one of the interlocutors. That is, we have to decide from incomplete information, guessing the intentions of the silent person. Implicatures allow to make inferences from what is said, but we can also infer from omission, or specifically from intentional silence in a conversation. In some contexts, not saying
Overlapping clustering algorithms have shown to be effective for clustering documents. However, the current overlapping document clustering algorithms produce a big number of clusters, which make them little useful for the user. Therefore, in this paper, we propose a k-means based method for overlapping document clustering, which allows to specify by the user the number of groups to be built. Our experiments with different corpora show that our proposal allows obtaining better results in terms of FBcubed than other recent works for overlapping document clustering reported in the literature.
Document clustering has become an important task for processing the big amount of textual information available on the Internet. On the other hand, k-means is the most widely used algorithm for clustering, mainly due to its simplicity and effectiveness. However, k-means becomes slow for large and high dimensional datasets, such as document collections. Recently the FPAC algorithm was proposed to mitigate this problem, but the improvement in the speed was reached at the cost of reducing the quality of the clustering results. For this reason, in this paper, we introduce an improved FPAC algorithm, which, according our experiments on different document collections, allows obtaining better clustering results than FPAC, without highly increasing the runtime.
Irony detection is a not trivial problem and can help to improve natural language processing tasks as sentiment analysis. When dealing with social media data in real scenarios, an important issue to address is data skew, i.e. the imbalance between available ironic and non-ironic samples available. In this work, the main objective is to address irony detection in Twitter considering various degrees of imbalanced distribution between classes. We rely on the emotIDM irony detection model. We evaluated it against both benchmark corpora and skewed Twitter datasets collected to simulate a realistic distribution of ironic tweets. We carry out a set of classification experiments aimed to determine the impact of class imbalance on detecting irony, and we evaluate the performance of irony detection when different scenarios are considered. We experiment with a set of classifiers applying class imbalance techniques to compensate class distribution. Our results indicate that by using such techniques, it is possible to improve the performance of irony detection in imbalanced class scenarios.
This paper describes our proposal for Sentiment Analysis in Twitter for the Spanish language. The main characteristics of the system are the use of word embedding specifically trained from tweets in Spanish and the use of self-attention mechanisms that allow to consider sequences without using convolutional nor recurrent layers. These self-attention mechanisms are based on the encoders of the Transformer model. The results obtained on the Task 1 of the TASS 2019 workshop, for all the Spanish variants proposed, support the correctness and adequacy of our proposal.
This paper proposes a sentiment analysis framework based on ranking learning. The framework utilizes BERT model pre-trained on large-scale corpora to extract text features and has two sub-networks for different sentiment analysis tasks. The first sub-network of the framework consists of multiple fully connected layers and intermediate rectified linear units. The main purpose of this sub-network is to learn the presence or absence of various emotions using the extracted text information, and the supervision signal comes from the cross entropy loss function. The other sub-network is a ListNet. Its main purpose is to learn a distribution that approximates the real distribution of different emotions using the correlation between them. Afterwards the predicted distribution can be used to sort the importance of emotions. The two sub-networks of the framework are trained together and can contribute to each other to avoid the deviation from a single network. The framework proposed in this paper has been tested on multiple datasets and the results have shown the proposed framework’s potential.
In this work we experiment with the hypothesis that words subjects use can be used to predict their psychological attachment style (secure, fearful, dismissing, preoccupied) as defined by Bartholomew and Horowitz. In order to verify this hypothesis, we collected a series of autobiographic texts written by a set of 202 participants. Additionally, a psychological instrument (Frías questionnaire) was applied to these same participants to measure their attachment style. We identified characteristic patterns for each style of attachment by means of two approaches: (1) mapping words into a word space model composed of unigrams, bigrams and/or trigrams on which different classifiers were trained (Naïve Bayes (NB), Bernoulli NB, Multinomial NB, Multilayer Perceptrons); and (2) using a word-embedding based representation and a neural network architecture based on different units (LSTM, Gated Recurrent Units (GRU) and Bilateral GRUs). We obtained the best accuracy of 0.4079 for the first approach by using a Boolean Multinomial NB on unigrams, bigrams and trigrams altogether, and an accuracy of 0.4031 for the second approach using Bilateral GRUs.
In recent times, sentiment analysis research has achieved tremendous impetus on English textual data, however, a very less amount of research has been focused on Nepali textual data. This work is focused towards Nepali textual data. We have explored machine learning approaches and proposed a lexicon-based approach using linguistic features and lexical resources to perform sentiment analysis for tweets written in Nepali language. This lexicon-based approach, first pre-process the tweet, locate the opinion-oriented features and then compute the sentiment polarity of tweet. We have investigated both conventional machine learning models (Multinomial Naïve Bayes (NB), Decision Tree, Support Vector Machine (SVM) and logistic regression) and deep learning models (Convolution Neural Network (CNN), Long Short-Term Memory (LSTM) and CNN-LSTM) for sentiment analysis of Nepali text. These machine learning models and lexicon-based approach have been evaluated on tweet dataset related to Nepal Earthquake 2015 and Nepal blockade 2015. Lexicon based approach has outperformed than conventional machine learning models. Deep learning models have outperformed than conventional machine learning models and lexicon-based approach. We have also created Nepali SentiWordNet and Nepali SenticNet sentiment lexicon from existing English language resources as by-product.
Poem is a spontaneous flow of emotions. There are several emotion detection systems to identify emotions from speech, gestures, and text (blogs, newspapers, stories and medical reports). Since such systems do not exist for poetry, we take the first step in building a system to recognize emotions in poetry by constructing a benchmark corpus, the PERC (
Creation of dictionaries of abstract and concrete words is a well-known task. Such dictionaries are important in several applications of text analysis and computational linguistics. Usually, the process of assembling of concreteness scores for words begins with a lot of manual work. However, the process can be automated significantly using information from large corpora. In this paper we combine two datasets: a dictionary with concreteness scores of 40,000 English words and the GoogleBooks Ngram dataset, in order to test the following hypothesis: in text concrete words tend to occur with more concrete words, than with abstract words (and inverse: abstract words tend to occur with more abstract words, than with concrete words). Using the hypothesis, we proposed a method for automatic evaluation concreteness scores of words using a small amount of initial markup.
Passage retrieval is an important stage of question answering systems. Closed domain passage retrieval, e.g. biomedical passage retrieval presents additional challenges such as specialized terminology, more complex and elaborated queries, scarcity in the amount of available data, among others. However, closed domains also offer some advantages such as the availability of specialized structured information sources, e.g. ontologies and thesauri, that could be used to improve retrieval performance. This paper presents a novel approach for biomedical passage retrieval which is able to combine different information sources using a similarity matrix fusion strategy based on convolutional neural network architecture. The method was evaluated over the standard BioASQ dataset, a dataset specialized on biomedical question answering. The results show that the method is an effective strategy for biomedical passage retrieval able to outperform other state-of-the-art methods in this domain.
RDF self-indexes compress the RDF collection and provide efficient access to the data without a previous decompression (via the so-called SPARQL triple patterns). HDT is one of the reference solutions in this scenario, with several applications to lower the barrier of both publication and consumption of Big Semantic Data. However, the simple design of HDT takes a compromise position between compression effectiveness and retrieval speed. In particular, it supports scan and subject-based queries, but it requires additional indexes to resolve predicate and object-based SPARQL triple patterns. A recent variant,
Currently, the semantic analysis is used by different fields, such as information retrieval, the biomedical domain, and natural language processing. The primary focus of this research work is on using semantic methods, the cosine similarity algorithm, and fuzzy logic to improve the matching of documents. The algorithms were applied to plain texts in this case CVs (resumes) and job descriptions. Synsets of WordNet were used to enrich the semantic similarity methods such as the Wu-Palmer Similarity (WUP), Leacock-Chodorow similarity (LCH), and path similarity (hypernym/hyponym). Additionally, keyword extraction was used to create a postings list where keywords were weighted. The task of recruiting new personnel in the companies that publish job descriptions and reciprocally finding a company when workers publish their resumes is discussed in this research work. The creation of a new gold standard was required to achieve a comparison of the proposed methods. A web application was designed to match the documents manually, creating the new gold standard. Thereby the new gold standard confirming benefits of enriching the cosine algorithm semantically. Finally, the results were compared with the new gold standard to check the efficiency of the new methods proposed. The measures used for the analysis were precision, recall, and f-measure, concluding that the cosine similarity weighted semantically can be used to get better similarity scores.
Email is one of the most popular ways of communication. Nevertheless, it is also a potential tool to deceive and fill users with unwanted publicity, which reduces productivity. To alleviate such fact, a common solution has been building machine learning models based on the content of emails to automatically separate emails (spam vs ham). In this work, a study of a set of machine learning models and content-based features for the problem of cross-dataset email classification is presented. This problem consists in training and testing the models using different datasets; considering the fact that the datasets were collected under different independent setups. This has the purpose of simulating future variable or unpredictable conditions in the emails content distributions as could happen in a real setting, where models are trained using emails from a certain period of time, group of users or accounts, but tested with emails from other users or accounts. Experiments were conducted with the models and features using different datasets and two setups, same-dataset, and cross-dataset, to show the complexity of the later. The performance was evaluated using the Area Under the ROC Curve, a common metric in email classification. The results show interesting insights for the problem.
Natural Ontologies are presented in this work as a useful tool to model the way in which concepts are organized inside the human mind. In order to be compared, ontologies are represented as matrices and an elastic matching technique is used. For this purpose, a distance measure called Modern Fréchet is proposed, which is an approximation to the NP-Complete problem of elastic matching between matrices. An applied case of study is presented in which human knowledge is compared among different groups of people in the Computer Science domain.
This work presents a method for data gathering to construct a corpus related to speech disorders in children; such corpus will serve as the base to generate some semi-automatic ontologies, in order to become a computational model to support therapists for diagnosis and possible treatment. Speech disorders, phonemes and some additional information are classified using taxonomies obtained from speech disorders specialized literature. Based on the obtained taxonomies, the ontologies, which structure and formalize concepts defined by the main topic authors, are developed. The ontologies are constructed following some parts of classic methodologies and their subsequent validation is made through competency questions. The development of the model is based on Natural Language Processing (NLP) and Information Retrieval (IR) techniques. Integration of the ontologies is made to be able to make a classification based in problematic phonemes; this is suggested as a complement to the diagnostic tool in the model.
In this paper, we introduce a Platform for Non-Intrusive Assistance (named PIANI), as an assistance platform for elderly people able to do activities in outdoor environments without strict supervision. PIANI includes an ontology used to characterize outdoor activities of interest (activities to be observed). PIANI also defines a risk level of the activity that an elderly person is currently doing out of his home by comparing such activity to its characterization. In addition, the proposed platform uses the smartphone of the person in order to collect geographic and time information, which is used by PIANI to infer activity risk and send alert notifications based on semantic knowledge base. An experimental test was developed as a proof of concept about the utilization of PIANI to identify outdoors activities of elderly people, compute a level of risk and finally send non intrusive alert notification to the user.
In the context of digital social media, where users have multiple ways to obtain information, it is important to have tools to detect the authorship within a corpus supposedly created by a single author. With the tremendous amount of information coming from social networks there is a lot of research concerning author profiling, but there is a lack of research about the authorship identification. In order to detect the author of a group of tweets, a Naïve Bayes classifier is proposed which is an automatic algorithm based on Bayes’ theorem. The main objective is to determine if a particular tweet was made by a specific user or not, based on its content. The data used correspond to a simple data set, obtained with the Twitter API, composed of four political accounts accompanied by their username and tweet identifier as it is mixed with multiple user tweets. To describe the performance of the classification model and interpret the obtained results, a confusion matrix is used as it contains values like accuracy, sensitivity, specificity, Kappa measure, the positive predictive and negative predictive value. These results show that the prediction model, after several cases of use, have acceptable values against the observed probabilities.
Interest has grown around the classification of stance that users assume within online debates in recent years. Stance has been usually addressed by considering users posts in isolation, while social studies highlight that social communities may contribute to influence users’ opinion. Furthermore, stance should be studied in a diachronic perspective, since it could help to shed light on users’ opinion shift dynamics that can be recorded during the debate. We analyzed the political discussion in UK about the BREXIT referendum on Twitter, proposing a novel approach and annotation schema for stance detection, with the main aim of investigating the role of features related to social network community and diachronic stance evolution. Classification experiments show that such features provide very useful clues for detecting stance.
The aim of the author profiling task is to automatically predict various traits of an author (e.g. age, gender, etc.) from written text. The problem of author profiling has been mainly treated as a supervised text classification task. Initially, traditional machine learning algorithms were used by the researchers to address the problem of author profiling. However, in recent years, deep learning has emerged as a state-of-the-art method for a range of classification problems related to image, audio, video, and text. No previous study has carried out a detailed comparison of deep learning methods to identify which method(s) are most suitable for same-genre and cross-genre author profiling. To fulfill this gap, the main aim of this study is to carry out an in-depth and detailed comparison of state-of-the-art deep learning methods, i.e. CNN, Bi-LSTM, GRU, and CRNN along with proposed ensemble methods, on four PAN Author Profiling corpora. PAN 2015 corpus, PAN 2017 corpus and PAN 2018 Author Profiling corpus were used for same-genre author profiling whereas PAN 2016 Author Profiling corpus was used for cross-genre author profiling. Our extensive experimentation showed that for same-genre author profiling, our proposed ensemble methods produced best results for gender identification task whereas CNN model performed best for age identification task. For cross-genre author profiling, the GRU model outperformed all other approaches for both age and gender.
Engaged customers are a very import part of current social media marketing. Public figures and brands have to be very careful about what they post online. That is why the need for accurate strategies for anticipating the impact of a post written for an online audience is critical to any public brand. Therefore, in this paper, we propose a method to predict the impact of a given post by accounting for the content, style, and behavioral attributes as well as metadata information. For validating our method we collected Facebook posts from 10 public pages, we performed experiments with almost 14000 posts and found that the content and the behavioral attributes from posts provide relevant information to our prediction model.
The task of author profiling aims to distinguish the author’s profile traits from a given content. It has got potential applications in marketing, forensic analysis, fake profile detection, etc. In recent years, the usage of bi-lingual text has raised due to the global reach of social media tools as people prefer to use language that expresses their true feelings during online conversations and assessments. It has likewise impacted the use of bi-lingual (English and Roman-Urdu) text in the sub-continent (Pakistan, India, and Bangladesh) over social media. To develop and evaluate methods for bi-lingual author profiling, benchmark corpora are needed. The majority of previous efforts have focused on developing mono-lingual author profiling corpora for English and other languages. To fulfill this gap, this study aims to explore the problem of author profiling on bi-lingual data and presents a benchmark corpus of bi-lingual (English and Roman-Urdu) tweets. Our proposed corpus contains 339 author profiles and each profile is annotated with six different traits including age, gender, education level, province, language, and political party. As a secondary contribution, a range of deep learning methods, CNN, LSTM, Bi-LSTM, and GRU, are applied and compared on the three different bi-lingual corpora for age and gender identification, including our proposed corpus. Our extensive experimentation showed that the best results for both gender identification task
In many areas of professional development, the categorization of textual objects is of critical importance. A prominent example is the attribution of authorship, where symbolic information is manipulated using natural language processing techniques. In this context, one of the main limitations is the necessity of a large number of pre-labeled instances for each author that is to be identified. This paper proposes a method based on the use of n-grams of characters and the use of the web to enrich the training sets. The proposed method considers the automatic extraction of the unlabeled examples from the Web and its iterative integration into the training data set. The evaluation of the proposed approach was done by using a corpus formed by poems corresponding to 5 contemporary Mexican poets. The results presented allow evaluating the impact of the incorporation of new information into the training set, as well as the role played by the selection of classification attributes using information gain.
The task of Extractive Multi-Document Text Summarization (EMDTS) aims at building a short summary with essential information from a collection of documents. In this paper, we propose an EMDTS method using a Genetic Algorithm (GA). The fitness function considering two unsupervised text features: sentence position and coverage. We propose the binary coding representation, selection, crossover, and mutation operators. We test the proposed method on the DUC01 and DUC02 data set, four different tasks (summary lengths 200 and 400 words), for each of the collections of documents (in total, 876 documents) are tested. Besides, we analyze the most frequently used methodologies to summarization. Moreover, different heuristics such as topline, baseline, baseline-random, and lead baseline are calculated. In the results, the proposed method achieves to improve the state-of-art results.
In this paper, we present an extractive approach to document summarization, the Siamese Hierarchical Transformer Encoders system, that is based on the use of siamese neural networks and the transformer encoders which are extended in a hierarchical way. The system, trained for binary classification, is able to assign attention scores to each sentence in the document. These scores are used to select the most relevant sentences to build the summary. The main novelty of our proposal is the use of self-attention mechanisms at sentence level for document summarization, instead of using only attentions at word level. The experimentation carried out using the CNN/DailyMail summarization corpus shows promising results in-line with the state-of-the-art.
The methods of Automatic Extractive Summarization (AES) uses the features of the sentences of the original text to extract the most important information that will be considered in summary. It is known that the first sentences of the text are more relevant than the rest of the text (this heuristic is called baseline), so the position of the sentence (in reverse order) is used to determine its relevance, which means that the last sentences have practically no possibility of being selected. In this paper, we present a way to soften the importance of sentences according to the position. The comprehensive tests were done on one of the best AES methods using the bag of words and n-grams models with the with DUC02 and DUC01 data sets to determine the importance of sentences.
A gesture elicitation study consists of a popular method for eliciting a sample of end end users to propose gestures for executing functions in a certain context of use, specified by its users and their functions, the device or the platform used, and the physical environment in which they are working. Gestures proposed in such a study needs to be classified and, perhaps, extended in order to feed a gesture recognizer. To support this process, we conducted a full-body gesture elicitation study for executing functions in a smart home environment by domestic end users in front of a camera. Instead of defining functions opportunistically, we define them based on a taxonomy of abstract tasks. From these elicited gestures, a XML-compliant grammar for specifying resulting gestures is defined, created, and implemented to graphically represent, label, characterize, and formally present such full-body gestures. The formal notation for specifying such gestures is also useful to generate variations of elicited gestures to be applied on-the-fly on gestures in order to allow one-shot learning.
Urdu is the most popular language in Pakistan which is spoken by millions of people across the globe. While English is considered the dominant web content language, characteristics of Urdu language web content are still unknown. In this paper, we study the World-Wide-Web (WWW) by focusing on the content present in the Perso-Arabic script. Leveraging from the Common Crawl Corpus, which is the largest publicly available web content of 2.87 billion documents for the period of December 2016, we examine different aspects of Urdu web content. We use the Compact Language Detector (CLD2) for language detection. We find that the global WWW population has a share of 0.04% for Urdu web content with respect to document frequency. 70.9% of the top-level Urdu domains consist of .
The paper presents a new corpus for fake news detection in the Urdu language along with the baseline classification and its evaluation. With the escalating use of the Internet worldwide and substantially increasing impact produced by the availability of ambiguous information, the challenge to quickly identify fake news in digital media in various languages becomes more acute. We provide a manually assembled and verified dataset containing 900 news articles, 500 annotated as real and 400, as fake, allowing the investigation of automated fake news detection approaches in Urdu. The news articles in the truthful subset come from legitimate news sources, and their validity has been manually verified. In the fake subset, the known difficulty of finding fake news was solved by hiring professional journalists native in Urdu who were instructed to intentionally write deceptive news articles. The dataset contains 5 different topics: (i) Business, (ii) Health, (iii) Showbiz, (iv) Sports, and (v) Technology. To establish our Urdu dataset as a benchmark, we performed baseline classification. We crafted a variety of text representation feature sets including word
Classification of research articles into different subject areas is an extremely important task in bibliometric analysis and information retrieval. There are primarily two kinds of subject classification approaches used in different academic databases: journal-based (aka source-level) and article-based (aka publication-level). The two popular academic databases- Web of Science and Scopus- use journal-based subject classification scheme for articles, which assigns articles into a subject based on the subject category assigned to the journal in which they are published. On the other hand, the recently introduced Dimensions database is the first large academic database that uses article-based subject classification scheme that assigns the article to a subject category based on its contents. Though the subject classification schemes of Web of Science have been compared in several studies, no research studies have been done on comparison of the article-based and journal-based subject classification systems in different academic databases. This paper aims to compare the accuracy of subject classification system of the three popular academic databases: Web of Science, Scopus and Dimensions through a large-scale user-based study. Results show that the commonly held belief of superiority of article-based subject classification over the journal-based subject classification scheme does not hold at least at the moment, as Web of Science appears to have the most accurate subject classification.
With evolution of knowledge disciplines and cross fertilization of ideas, research outputs reported as scientific papers are now becoming more and more interdisciplinary. An interdisciplinary research work usually involves ideas and approaches from multiple disciplines of knowledge applied to solve a specific problem. In many cases the interdisciplinary areas eventually emerge as full-fledged disciplines. In the last two decades, several approaches have been proposed to measure the Interdisciplinarity of a scientific article, such as propositions based on authorship, references, set of keywords etc. Among all these approaches, reference-set based approach is most widely used. The diversity of knowledge in the reference set has been measured with three parameters, namely
Deep learning architectures based on self-attention have recently achieved and surpassed state of the art results in the task of unsupervised aspect extraction and topic modeling. While models such as neural attention-based aspect extraction (ABAE) have been successfully applied to user-generated texts, they are less coherent when applied to traditional data sources such as news articles and newsgroup documents. In this work, we introduce a simple approach based on sentence filtering in order to improve topical aspects learned from newsgroups-based content without modifying the basic mechanism of ABAE. We train a probabilistic classifier to distinguish between out-of-domain texts (outer dataset) and in-domain texts (target dataset). Then, during data preparation we filter out sentences that have a low probability of being in-domain and train the neural model on the remaining sentences. The positive effect of sentence filtering on topic coherence is demonstrated in comparison to aspect extraction models trained on unfiltered texts.
LinkedIn is a social medium oriented to professional career handling and networking. In it, users write a textual profile on their experience, and add skill labels in a free format. Users are able to apply for different jobs, but specific feedback on the appropriateness of their application according to their skills is not provided to them. In this work we particularly focus on applicants of the project management branch from information technologies—although the presented methodology could be extended to any area following the same mechanism. Using the information users provide in their profile, it is possible to establish the corresponding level in a predefined Project Manager career path (PM level). 1500+ experiences and skills from 300 profiles were manually tagged to train and test a model to automatically estimate the PM level. In this proposal we were able to perform such prediction with a precision of 98%. Additionally, the proposed model is able to provide feedback to users by offering a guideline of necessary skills to be learned to fulfill the current PM level, or those needed in order to upgrade to the following PM level. This is achieved through the clustering of skill qualification labels. Results of experiments with several clustering algorithms are provided as part of this work.
There is a lot of cultural heritage information in historical documents that have not been explored or exploited yet. Lower-Baseline Localization (LBL) is the first step in information retrieval from images of manuscripts where groups of handwritten text lines representing a message are identified. An LBL method is described depending on how the features of the writing style of an author are treated: the character shape and size, gap between characters and between lines, the shape of ascendant and descendant strokes, character body, space between characters, words and columns, and touching and overlapping lines. For example, most of the supervised LBL methods only analyze the gap between characters as part of the preprocessing phase of the document and the rest of features of the writing style of the author are left for the learning phase of the classifier. For such reason, supervised LBL methods tend to learn particular styles and collections. This paper presents an unsupervised LBL method that explicit analyses all the features of the writing style of the author and processes the document by windows. In this sense, the proposed method is more independent from the writing style of the author, and it is more reliable with new collections in real scenarios. According to the experimentation, the proposed method surpasses the state-of-the-art methods with the standard READ-BAD historical collection with 2,036 manuscripts and 132,124 manually annotated baselines from 9 libraries in 500 years.
This paper presents a novel deep learning based approach to solving
We tried to determine if emotive self-feedback from conscious assessment of artists’ own works generates sufficient impetus for accomplishment of goals. Self-reports from participants of an ‘experimental’ group working independently and
Currently, there is a great necessity in the organizations to support communication, collaboration, and coordination — important aspects that characterize collaborative systems—between its workers and enterprises; in order to simplify and improve their production processes. However, the development and maintenance of these systems are very complex. Although several proposals to develop them have been made, they usually lack theoretical models, which allow specifying and creating both group and interactive activities in a conceptual and formal way to sustain the requirements of group work. Therefore, this paper PRoposes an Ontological Model for developing collaboratIve SystEms (PROMISE), it tries to be a guide for the analysis, design, and implementation of such systems in a formal, explicit manner. This model is based on an ontology, created using OWL (Web Ontology Language), providing a model of knowledge about in what way entities should be used and combined to control the execution of a set of orderly steps to develop these systems. Furthermore, this ontology has been validated through a set of academic’s projects, showing be great usefulness to developers.
Introduction Peripheral arterial disease (PAD) is a fairly common degenerative vascular condition in diabetic patients that leads to inadequate blood flow (BF), this disease is mainly due to atherosclerosis that causes chronic narrowing of arteries, which can precipitate acute thrombotic events. In patients with diabetes, atherosclerosis is the main reason for reducing life expectancy, as long as diabetic nephropathy and retinopathy are the largest contributors to end-stage renal disease and blindness, respectively. Objective This was an assessment of dilatation of the blood vessels on diabetic patients vs. healthy volunteers by using digital processing of imaging’s. Materials and Methods The study subject was ultrasound imaging processing of blood vessels dilation on low extremities of diabetic patients, the results were compared with ultrasound images of healthy subjects. Results The digital images processing suggests that there is a significant difference among images experimental of the diabetic group and healthy volunteers’ images, the control group. Discussion The digital imaging processing performed in the Matlab platform is an adequate procedure for blood vessels dilation analysis of the ultrasound images taken from the lower extremities in diabetic patients.