Abstract
The qualitative analysis results of teachers’ abilities are difficult to quantify, and ability problems in the teaching process are difficult to be effectively measured. In order to study methods to improve teachers’ teaching abilities, this paper builds a corresponding teacher competence evaluation model based on machine learning and digital twin technology, establishes a data collection model for teachers’ professional competence, and establishes a data fusion model. It includes data cleaning model based on XML information template, data integration model, multi-index screening mechanism and clustering strategy based on perturbation attributes. On this basis, this paper uses decision tree algorithm, random forest algorithm and neural network algorithm to construct three scheduling rule mining models aiming at teachers’ professional ability. In addition, this paper establishes a digital twin-driven multi-knowledge model scheduling optimization architecture that uses the three scheduling rules mined. The research results show that the model constructed in this paper has good performance.
Introduction
In the context of the continuous development and reform of my country’s higher education and its modernization, colleges and universities continue to expand enrollment to provide manpower support and intellectual contributions for the development of the new era. However, although higher education is booming, enrollment expansion has also brought some negative effects. Among them, the most important issue is that the speed of improving the quality of Chinese higher education does not match the speed of the number of people climbing. Teachers’ teaching ability is an important part of education elements and a key factor affecting the quality of education. As the main force of the university system, the development of teachers’ teaching ability will have a great impact on the future development of higher education. At the same time, regardless of the construction of double first-class or university rankings, the development level of teachers’ teaching ability is regarded as one of the core elements of evaluation. In this way, improving the teaching development ability of teachers is very important for improving the development of higher education and the quality of talent cultivation. Moreover, it also plays an important role in promoting the modernization of education development, improving the competitiveness and international influence of Chinese universities [1].
At present, the teaching ability of teachers in colleges and universities in my country is relatively inadequate. On the one hand, they generally lack the sense of autonomy to improve teaching ability, and even do not fully participate in teaching, and they still position themselves as passive roles, such as “teacher”, classroom monologues, and course performers [2]. Moreover, their teaching concepts are not compatible with the requirements of modern education, and they are not aware of the need to improve their teaching abilities. On the other hand, current college students are eagerly looking forward to obtaining guidance in learning professional knowledge and skills and career planning guidance. However, teachers are shortly employed and generally lack teaching experience, and most college teachers have not received training in teacher professional education before entering the university system, and only master the silent knowledge of the theoretical basis of the subject. In addition, they also do not know how to use reasonable methods and means to apply theory to teaching practice. Therefore, there is an urgent need to provide effective methods externally to help them transform internalized knowledge into teaching content. On this basis, my country has continuously strengthened the importance of teachers’ teaching abilities in recent years. The Ministry of Education emphasizes the need to implement the extensive establishment of the Teacher Development Center to provide a strong guarantee for the development of teachers’ teaching abilities. However, at present, from the government to colleges and universities, there is still a general lack of clear awareness of the development of teachers’ teaching ability, and a long-term stable operating mechanism has not been formed. How to train and develop the teaching level of college teachers in the current dilemma still needs to be more standardized. Therefore, it is particularly important to carry out research to implement relevant national guidance and policies. Nowadays, the teaching ability of college teachers has become a research hotspot in Chinese academic circles. However, the empirical research based on the actual situation and actual needs of the teacher group’s teaching ability improvement is still insufficient, and it is difficult to guide us to conduct more in-depth research on the actual path of improvement [3].
Therefore, to explore the current needs and current status of teachers’ teaching ability improvement, to analyze the problems and obstacles from a theoretical perspective, and to propose a breakthrough path and its focus can effectively improve the quality of talent training, conform to modern education, and promote the growth and development of college teachers. Moreover, it is conducive to enrich and develop the theories related to the teaching abilities of college teachers. The study of its teaching ability is one of the important subjects of teacher education research. At present, most of the relevant theoretical research is focused on the professional development of teachers in the stage of compulsory education, but there is very little theoretical research on improving the ability of college teachers, especially in teaching. Therefore, it is possible to strengthen the grasp of teachers’ teaching development laws by deepening the basic understanding of their teaching development needs and processes.
Related work
The literature [4] believed that teachers should put the independent development of students in the first place and put students in the center of teaching. Under the guidance of this viewpoint, the literature further concluded that teaching research has a great role in promoting student learning and teacher teaching, which lays the foundation for the vigorous development of teaching ability research. After that, more scholars began to pay attention to the role of teaching ability in teaching activities. Among them, literature [5] believed that teaching ability needs special training, and college teachers should be good at synthesizing multiple teaching methods, exert teaching wisdom, and enhance teaching effectiveness. The literature [6] proposed that the academic dimensions are exploration, synthesis, application and teaching, and proposed teaching academic theories. This theory breaks through the teaching boundary and to a certain extent, regards teaching as the inferior concept of academia, endows teaching with research functions, and finds a certain balance between teaching and research. This view has been strongly recognized by colleges and universities, making teaching from closed and single to open and pluralistic, and it also provides good theoretical support for young teachers to invest in teaching research.
When carrying out evaluation research on teaching ability, literature [7] found that teaching ability should involve aspects such as teaching plan, inheritance theory basis and implementation of teaching activities. After conducting in-depth research, the literature [8] proposed that teaching ability includes many aspects. For example, teachers must have a serious attitude and certain learning ability in dealing with knowledge, give them care and attention in dealing with students, provide a relaxed teaching atmosphere, understand students and cherish students, and be good at guiding and asking questions. The literature [9] put forward that the teaching guidance ability, knowledge construction ability and teaching research ability of teachers in professional development are key. The literature [10] proposed that the core elements of effective teaching are teaching evaluation ability, ability to provide knowledge and self-renewal ability. The literature [11] believed that the teaching development of college teachers is related to the individual, environment and activities. Among them, personal factors include development awareness, career planning, etc., the environment includes school teaching culture, departments and families, etc., and activities include training and academic exchanges related to ability enhancement. The literature [12] divided the influencing factors into two aspects: individual and organization, which includes more and more detailed reasons. Personal factors such as stress, age, personality, etc. affect teaching ability from within, while organizational environment, academic community, school system, etc. affect teaching ability from outside.
The literature [13] conducted an in-depth study on ways to improve teaching ability, and concluded that the instructional activity of teaching ability is an effective method, which is widely favored by young teachers. Teaching guidance is usually conducted by senior teaching experts, and indicates the development direction for young teachers’ teaching content, skills, and research directions. This close contact allows young teachers to fully observe the teaching process of teaching experts, and enables young teachers to acquire tacit knowledge in the process of timely communication. In terms of the form of teaching guidance, the literature [14] believed that the network can better implement the guidance function. Moreover, the literature believed that the Internet can gather young teachers in an invisible group to give them a sense of belonging to a certain extent, generate a sense of cooperation, jointly solve teaching problems, and develop teaching abilities. Western countries with well-developed education such as the United States, Germany, and Canada pay attention to the research on the ability of college teachers, especially the teaching ability, which has a remarkable effect in improving the path. The literature [15] first proposed the establishment of a teaching development center and took the teaching development center as the main organization to develop teaching ability. In addition, we can improve teachers’ teaching ability through short-term training, exchange seminars, academic conferences, and expert lectures. Focusing on the special features of young teachers, Germany created a teaching ability training institution for this group to provide theoretical guidance and practical opportunities to promote its teaching ability [16]. The Higher Education Research and Development Center is the main institution for teacher training in Finland. Its characteristic is that it divides the development of teachers into different stages and trains them according to the characteristics of the stages. At the same time, the state takes teaching level and training achievements as the focus of assessment and encourages teachers to consciously join the training activities of the research center [17]. The training puts the teacher’s subjectivity at the center, and the form of activities is flexible, and there is a high degree of flexibility in terms of time and place. Because teachers can choose according to their own circumstances, training is welcomed by college teachers [18]. The United Kingdom is committed to forming a joint training force, organizing regional universities to form training alliances, and promoting the improvement of teaching ability through a variety of activities. Moreover, it promotes peer-to-peer communication among college teachers while conducting teaching training to help teachers complete teaching tasks and develop educational wisdom. This project has received positive feedback and appreciation from college teachers [19]. The literature [24] addresses the various problems in the field of vehicle communication with the suggestion of a mutual unified and dispersed spectrum sensing model. The application of the mutual cognitive paradigm minimizes conflict and multiple unknown problems. The literature [25] discusses the problem of vast volumes of big data and introduces the SmartBuddy idea of an adaptive and smart world incorporating human activity and human dynamics. The literature [26] talks about the development in parallel reconfigurable computing systems of a directed acyclic graph for video coding algorithms for motion estimation. Partitioning algorithm also plays a major role in speeding up the production of images. The article [27] deals with leveraging IoT and BigData Analytics in real-time applications using the Hadoop platform. The above-mentioned processes enable the deployment of an IoT-based Smart City. The article [28] centers on IoT and its major part in sophisticating the human practices and endeavors. This paper moreover managed with the collection of different information from different assets that are associated to the web [29, 30].
Design of scheduling rule mining model based on improved random forest algorithm
Bagging is a commonly used method based on integrated learning ideas to improve the accuracy of learning algorithms. A set of training instance sample set P ={ simple1, ⋯ , simple n } and a set of weak learning algorithm L ={ l1, ⋯ , l m } are given. First, the training instance samples are extracted from the training instance sample set with replacement to form m new training instance sample sets, and then these weak learning algorithms are used to obtain m weak learnersFinally, the final result is obtained by synthesizing the prediction results given by m learning in a certain way. If the problem to be solved is a classification problem, then the final result will be obtained by voting.
The random forest algorithm used in this paper adopts C4.5 decision tree {C (X, Θ q ) , q = 1, ⋯ , Q } as the basic classifier, and through bagging strategy and feature attribute random extraction strategy, it is further combined to form an integrated classifier. We assume that the set of characteristic attributes constructed by the C4.5 decision tree is {Θ q }, and the parameter q represents the number of decision trees in the constructed random forest. The construction process of the random forest is shown in Fig. 1. First, we randomly sample back from the data that was converted into the form of scheduling instances after completing the data integration to form q new training instance sets. These q new training instance sets are used, and each time m feature attributes are randomly selected from {Θ q }, and q new decision trees are obtained by using the construction method of C4.5 decision trees. Among them, the C4.5 decision tree is not pruned after construction, allowing the decision tree to grow freely. The resulting random forest can complete new classification tasks, and the classification results will be obtained by a mechanism of obeying the majority. The final classification decision formula is as follows [20]:

The construction process of random forest.
In the formula, C (X) represents a random forest algorithm, c i (X) represents a single C4.5 decision tree algorithm, and Y represents the classification result. Meanwhile, I represents an indicative function.
According to the law of large numbers, the random forest algorithm has an upper bound on the generalization error. The upper bound is as follows [21]:
The q in the formula represents the number of C4.5 decision trees in the constructed random forest. It can be seen from the formula that as the number of decision trees in the random forest increases, the generalization error of the random forest model will tend to upper bound. This shows that the random forest algorithm has strong generalization ability and can better avoid the problem of overfitting.
In order to meet the objective requirements of scheduling rule mining, this paper makes two improvements to the random forest algorithm.
First, a strategy to avoid similar decision trees is proposed. Through the random forest algorithm, the scheduling rule f is learned from the scheduling history related data. f is actually an estimate
In the formula, δ2 is noise, expressing the lower bound of expected generalization error that any learning algorithm can achieve in scheduling knowledge extraction.
The decision tree generated by the random forest algorithm through the Bagging (bootstrap aggregation) strategy has an approximate distribution. Therefore, the variance of the random forest algorithm can be regarded as the variance of a group of randomly distributed random variables. The variance calculation formula is shown in the following formula. It can be seen from the formula that for a random forest algorithm that needs to build multiple decision trees (n is larger), if the correlation ρ2 between decision trees can be reduced, the variance can be reduced, thereby improving the performance of the algorithm.
In the formula, n is the number of decision trees in the random forest, T i represents the i-th decision tree, and ρ represents the correlation between decision trees. At the same time, θ2 represents the variance of each decision tree.
Based on the above analysis, this paper proposes a strategy to avoid similar decision trees, which is used to reduce the correlation ρ between decision trees and to improve the performance of the random forest algorithm. If the similarity between the two decision numbers is greater than 60%, that is, they are considered to be similar decision trees, and the one that performs poorly on the classification of the test data needs to be eliminated. This strategy uses the similarity calculation method of decision trees proposed by Irene. The similarity between decision trees depends on the percentage of the same predictions that they use the same feature attributes and produce the same predictions for test cases. The similarity calculation formula is shown in formula (5).
In the formula, DT1 and DT2 represent the two decision trees for similarity calculation, k represents the same number of times that DT1 and DT2 classify the test instance, and r1nandr2n indicate the characteristic attributes used by DT1 and DT2 when the nth classification result is the same, c indicates the classification result. When r1n = r2n, that is, DT1 and DT2 use the same feature attributes to get the same classification result, I (r1n . c, r2n . c) = 1, otherwise it is 0. Nt is the number of test cases.
Second, the Bayesian voting mechanism is used. In the Bayesian voting mechanism, the calculation formula of the weighted voting result of each decision tree is as follows:
In the formula, WR represents the weighted result given by the decision tree, v represents the number of times the decision tree correctly classified the test instance, and m represents the number of times the test instance was incorrectly classified. Meanwhile, C represents the classification result given by this decision tree, and R represents the average value of the classification results given by all decision trees [23].
For a certain disturbed environment, the specific steps of using the improved random forest algorithm to mine the pseudo code of the scheduling rules are as follows:
Step 1: Random forest is constructed. First, the better scheduling data in the cluster corresponding to the disturbance environment is used as training data. On this basis, the training examples are extracted from the d2/d3 part of the better scheduling data with replacement to form q new training instance sets (bagging strategy). Finally, each training instance set randomly selects m attributes from d2/d3’s attribute set as feature attributes and calculates the best classification method to obtain q decision trees.
Step 2: The performance of decision tree classification is tested. The data of part d2/d3 of the better scheduling data that is not selected as the training instance is used as test data to test and record the classification performance of each decision tree.
Step 3: Similar decision tree strategy is avoided. The similarity between decision trees is calculated. If the similarity between two decision trees is greater than 60%, it is considered as a similar decision tree, and one of the poorer ones in the test performance needs to be eliminated.
Step 4: The weight of each decision tree is calculated. According to the performance of the classification of test data, the weight of each decision tree retained in the random forest is calculated, that is, v/(v + m)m/(v + m) in equation (6) is obtained.
The model of scheduling rules mining method based on random forest is shown in Fig. 2. After mining the random forest scheduling rules through the improved random forest algorithm mining, the random forest scheduling rules can be used to distinguish suitable machines or to obtain workpieces with high processing priority. For example, Operation X of a workpiece needs to select the most suitable machine in m1, m2, m3 for processing. First, the machine suitable for Operation X processing is found from m1 and m2, and the real-time information of m1 and m2 is input into the random forest scheduling rule 1 to obtain all decision tree classification results in this rule. These results include 1 and 2 (1 represents the former m1 is suitable, 2 represents the latter m2 is suitable). According to formula (6), the weighted voting results of each decision tree are obtained, and these results are averaged. If the average is less than 1.5, it indicates that the former m1 is more suitable for Operation X processing, otherwise m2 is a suitable machine. After obtaining the appropriate machines from m1 and m2, the same method is used to compare with m3, the machine most suitable for Operation X processing can be obtained.

Improved random forest algorithm mining model.
Figure 3 shows the structure of the GAP-RBF neural network. This is a GAP-RBF neural network with structure m - v - n. There are m neurons in the input layer, v neurons in the hidden layer, and n neurons in the output layer. X = (x1, ⋯ , x n ) T is the input vector offset, f (x) ={ f1 (x) , ⋯ , f n (x) } is the output vector offset, c i is the center value of the ith hidden layer neuron, and w is the output weight matrix.

RBF neural network structure.
The basis function of the neuron in the hidden layer of the GAP-RBF neural network takes the distance between the input vector and the center vector as the independent variable, and uses the radial basis function as the activation function. The commonly used radial basis functions are as follows. Among them, ∥* ∥ represents the Euclidean norm: Gaussian function (Gaussian function), the formula is as follows
Reflected sigmoidal function (abnormal S-shaped function), the formula is as follows
Cubic function, the formula is as follows
Thin plate spline interpolation function, the formula is as follows
Linear function, the formula is as follows
Multiple quadratic functions, the formula is as follows
Among the above six radial basis functions, the Gaussian function is the most commonly used one. In this paper, the Gaussian function is also selected as the activation function of the algorithm.
During the learning process of the GAP-RBF neural network, as the input samples are input, the neural network will accept the samples in order. Whenever the neural network accepts a sample, it will decide whether to add a new hidden layer neuron or delete the nearest neuron to the current input sample based on the Significance of the nearest hidden layer neuron to the input sample relative to the input sample, or adjust the parameters of hidden layer neurons. The Significance of the hidden neurons of the i-th is defined as follows:
In the formula, l represents the dimension of the input sample data, that is, the number of production attributes. X represents the variation range of the sample, and S (X) represents the range of the sample.
The hidden layer of the initial GAP-RBF neural network does not contain neurons. As the neural network accepts input samples one by one, the algorithm will activate new hidden layer neurons according to the growth criterion. The growth criteria for the smooth growth of hidden layer neurons are shown in (14), (15) and (16). Only when these three formulas are satisfied at the same time, the new hidden layer neurons will be activated.
In the formula, the distance ɛ n is the resolution, ɛ n = ɛmaxin the initial state, and then gradually decreases to ɛmin with the input of the sample. γ is the attenuation coefficient, emin is the expected accuracy of the network output, and k is the overlap factor. As k increases, the overlapping response area between hidden layer neurons also becomes larger.
When the growth criterion of smooth growth of hidden layer neurons is satisfied to activate new hidden layer neurons, The parameters of the new hidden layer neuron are as shown in formula (18), formula (19) and formula (20).
When the received input information does not meet the growth criterion, the parameter w = [w1, μ1, δ1, ⋯ , w
v
, μ
v
, δ
v
] of the hidden layer neuron needs to be adjusted by the EKF method, and the adjustment method is as follows:
In the above formula, k
n
is the Kalman gain vector, and its calculation formula is as follows:
In the above formula, R
n
is the noise variance, a
n
is the gradient vector, and a
n
is calculated by formula (23). At the same time,P
n
is the error covariance matrix, and P
n
is calculated by formula (24)
In the above formula, Q is a random step and I is the identity matrix. As the number of neurons in the hidden layer increases, the dimension of P
n
will also increase. The way it rises is as follows:
In the above formula, P0 is the initialization value. The dimension of unit array I is equal to the number of adjustable parameters that increase by activating a new hidden layer neuron.
As the neural network training process continues, the GAP-RBF algorithm will also delete some hidden layer neurons according to the reduction criteria. The reduction criteria are as follows
This paper takes GAP-RBF neural network as one of the algorithms for scheduling rule mining. The GAP-RBF neural network to be constructed is essentially a mapping of scheduling related information such as machine status information, workpiece status information, delivery information, etc., to scheduling decisions. The design process of GAP-RBF neural network is as follows:
(1) Network input and output are determined
The input to the network is part d2 and d3 of the scheduling-related data after data fusion has been completed.
The output of the network is that the vectory n is 1 or 2, which indicates the priority of the two machines and the workpiece compared with each other. 1 represents the former is suitable, 2 represents the latter is suitable.
(2) Selection of radial basis function
The radial basis function is the activation function of the hidden layer of the GAP-RBF neural network. Its function is to map the received input signal to the hidden layer neuron space through a linear transformation.
(3) Determination of core parameters of GAP-RBF neural network
For each disturbed environment after clustering, part d2 data and part d3 data are used as training data, respectively, to train the GAP-RBF neural network until the output error is less than a certain threshold, then the training is stopped. The connection weight, data center value and width value are obtained.
(4) Output function determination
After training to obtain the connection weight of the hidden layer neuron to the output layer neuron, the expression of the output function can be obtained, details as follows:
Based on the above GAP-RBF neural network design process, the GAP-RBF scheduling rule mining model shown in Fig. 4 is designed.

GAP-RBF scheduling rule mining model.
By using the GAP-RBF scheduling rule mining model in Fig. 4, the required scheduling rules can be mined. The resulting scheduling rule is the GAP-RBF scheduling rule obtained by training on scheduling-related data belonging to different disturbance environments. The usage method of the GAP-RBF scheduling rules obtained by mining is shown in Fig. 5. First of all, according to the current disturbance environment and the machine selection problem of the workpiece or the workpiece selection and processing problem of the idle machine, the GAP-RBF scheduling rule obtained from the data training in the corresponding cluster is found. After that, the required real-time information is input into the scheduling rules, and by the method of pairwise comparison, the machine selection problem of the workpiece and the workpiece selection processing problem of the idle machine can be solved.

Usage instruction of GAP-RBF scheduling rules.
To verify the performance of the model, this study first verified the operation of the model itself, mainly from the data response time and retrieval accuracy. 100 sets of data are set respectively, and the test results of the model are counted, and the statistical results are drawn into statistical diagrams. The first is the statistics of the response time of data processing. The results are shown in Table 1 and Fig. 6.
Statistical table of system data processing time
Statistical table of system data processing time

Statistical table of system data processing time.
From the above figures and tables, it can be seen that when this system performs a large amount of data processing, the processing time of each group of data does not exceed 100 ms, and the fastest processing time can reach 35 ms. This shows that the data processing speed of this system is faster. After that, this study conducted a data retrieval accuracy rate simulation, and the results obtained are shown in Table 2 and Fig. 7.
Statistical table of data retrieval accuracy

Statistical diagram of data retrieval accuracy.
It can be seen from Fig. 7 that the system data retrieval accuracy rate constructed in this article is basically above 99%, so it meets the actual needs.
On the basis of the above analysis, the model’s practical effect analysis is carried out, and it is applied to actual teaching, the actual teaching data is counted, and the model’s own data mining capabilities are used to discover the rules. The results obtained after teaching practice in a college are shown in Table 3 and Fig. 8.
Statistical table of model application results

Statistical diagram of model application results.
From the above analysis, we can see that when the system is applied to practical teaching, it can effectively identify the teaching ability of teachers and make targeted suggestions, so it has certain effects.
Teachers’ professional competence is the core of appraisal of teachers’ professional qualifications, and the promotion of teachers’ professional competence is the key to improving the quality of education, and it is the basic requirement for the construction of a strong country in higher education in my country. Moreover, its importance shows that it has extremely high research value. With the help of digital twin technology, this paper proposes a new digital twin-driven scheduling model. First, a data collection model for teachers’ professional competence is established, which includes a real-time data collection model containing adaptive data collection strategies and differentiated transmission strategies and an offline data collection model with the help of big data technology. At the same time, a data fusion model was established, which included a data cleaning model based on XML information templates, a data integration model, a multi-indicator screening mechanism, and a clustering strategy based on perturbed attributes. On this basis, this study uses C4.5 decision tree algorithm, random forest algorithm and GAP-RBF neural network algorithm to construct three scheduling rule mining models for teachers’ professional ability. Moreover, this research establishes a digital twin-driven multi-knowledge model scheduling optimization framework that uses the three scheduling rules that are mined. Finally, the model effect is verified. The results show that the practical effect of this model is good.
Footnotes
Acknowledgment
This paper was supported by (1) National Special Support Program for high-level Talents: Research and practice on four-stage teaching of “Yixian Xinhua Class”, Project number: 50000-31910018, PROJECT LEADER: WANG Tinghuai; (2) Guangdong Province Education Science “12TH FIVE-YEAR PLAN” Research Project: Research on the cultivation of independent learning ability of independent college students in the new media era, Project number: 2013JK340, PROJECT LEADER: WANG Donghui.
