Abstract
In this paper, we aim to solve the optimal tracking control problem for the Henon Mapping chaotic system using the direct heuristic dynamic programming (DHDP) setting with filtered tracking error. The fuzzy logic system is used to approximate the long-term utility function. Compared with the results for chaotic discrete-time system, the cost of the controller is reduced. The Lyapunov analysis approach is utilized to prove the stability of the chaotic system. It is shown that the tracking error, the adaptation law and the control input retain the property of uniformly ultimate boundedness. A simulation example is given to demonstrate the effectiveness of the proposed approach.
Keywords
1. Introduction
In the past decade, the control problem of the chaotic system has attracted much attention. Some traditional strategies have been designed for chaotic systems. In Wang and He (2008), to analyze the dynamical behavior of fractional order unified system, a projective synchronization approach of fractional order unified system was proposed based on the stability criterion of linear systems. A high precision fast projective synchronization method was proposed in Wang et al. (2009) through introducing an impact factor for chaotic (hyperchaotic) systems. An adaptive scheme based on Gaussian network was provided in Zhang et al. (1998) for regulation of uncertain chaotic systems. Subsequently, a lot of adaptive control methods were proposed one after the other in Hua et al. (2005), Chen et al. (2009), Chen and Chen (2009), Tang et al. (2013), Li (2012) and Li (2013a) for uncertain nonlinear chaotic systems. Specifically, a great amount of control approaches using the fuzzy logic systems and the neural networks for uncertain nonlinear systems had attracted much attention (Chen et al., 2013, 2014; Hu, 2009, Hu et al., 2008, Hu and Ma, 2008; Yang et al., 2008, 2009b; Li et al., 2010, 2011, 2012b, 2014a, 2014b, Li, 2012, 2013, 2014, Tong and Li, 2003, Tong et al., 2005, 2010, 2011).
The above results are obtained for the chaotic systems in continuous-time. In Yamamoto et al. (2001), a delayed feedback control was developed for chaotic discrete-time systems. Lu et al. (2001)'s study adaptive design using backstepping for a class of discrete-time chaotic systems with known or unknown parameters. However, the current methods do not consider the controller cost problem. It is significant work to find the optimal control performance.
Reinforcement learning has held great intuitive appeal and has attracted considerable attention in the past. But only recently has it made major advancements by implementing the temporal different (TD) learning method. The reinforcement learning based adaptive critic neural network (NN) approach has emerged as a promising tool to develop optimal NN controllers due to its potential to find approximate solutions to dynamic programming, where a strategic utility function (a long-term system performance measure) can be optimized. By contrast, dynamic programming (DP) provides truly optimal solutions to nonlinear stochastic dynamic systems. Heuristic dynamic programming (HPD) was proposed in the 1970s and the ideas were firmed up in the early 1990s under the names of adaptive critic designs. The original proposition for HDP was essentially the same as the formulation of reinforcement learning (RL) using TD methods. Specifically, this formulation falls exactly into the Bellman equation. HDP and the adaptive critics in general train a network to associate input states with action values. The direct heuristic dynamic programming (DHDP) method is an approximate DP, which was inspired by action-dependent HDP (ADHDP) and has been applied to large-scale complex realistic applications. Once again, the basic idea in adaptive critic design is to adapt the weights of the critic network to make the approximating utility function satisfy the modified Bellman equation.
In Yang and Jagannathan (2012), the adaptive reinforcement learning controller was proposed by using the neural networks for unknown nonlinear discrete-time multiple-input-multiple-output (MIMO) systems with the external disturbances. In the design, an action network that is designed to produce optimal signal and a critic network that evaluates the performance of the action network. In Lin (2003), Tang and Liu (2013), two adaptive reinforcement learning schemes were studied for robot manipulator with uncertainties. By designing the direct heuristic dynamic programming, a nonlinear tracking control setting with filtered tracking error was designed in Yang et al. (2009a). The stability analysis of the tracking system is given based on the Lyapunov stability. It is shown that the closed-loop tracking error and the approximating neural network weight estimates are uniformly ultimate boundedness.
The above methods are very good to solve optimal control problem. But the optimal control strategy is not obtained for the discrete-time chaotic systems. Thus, this paper tries to address an optimal control problem for Henon mapping chaos model with unknown parameters. The fuzzy logic systems are used to approximate the uncertainties of the systems. Based on direct heuristic dynamic programming Yang et al. (2009a), an optimal controller is designed compared with the method in Lu et al. (2001). The Lyapunov method is employed to analyze the stability of the closed-loop system. A simulation example is provided to verify the effectiveness of the proposed approach.
2. System description
The problem of controlling a class of discrete-time chaotic systems by using the DHDP approach will be studied. Consider the Henon Mapping chaos model
According to the chaos model (equation (1)), consider the chaotic system without the control input The chaotic phenomenon.
The tracking error
Given the desired trajectory
Rewriting equation (4) at k + 1, i.e.
3. Fuzzy approximation
The basic configuration of the fuzzy logic systems consists of fuzzifier, fuzzy rule base, fuzzy inference engine and defuzzifier. The fuzzy inference engine employs fuzzy rules to perform a mapping from an input linguistic vector
The fuzzy logic systems with the singleton fuzzifier, product inference engine and center-average defuzzifier are defined as follows
It has been proven that the fuzzy logic systems can uniformly approximate any a nonlinear discrete function to an arbitrary accuracy on a compact set. We introduce the following lemma. For any a given function Lemma 1
Given a compact set
In this paper, the control objective is to design an adaptive controller for the system (equation (1)) that all the signals in the closed-loop system remain uniformly ultimately bounded, the system state
4. Strategic utility function
Reinforcement signal
The goal of DHDP control is to optimize a long-term cost function
We use the fuzzy logic system to estimate the long-term cost function
By writing the Bellman equation into the prediction error
The objective function to be minimized based on the prediction error is defined as
By substituting equation (14) into the preceding equation, the weight-updating rule becomes
5. Controller design and stability analysis
Define control input
Assuming that the filtered tracking error system is stable, given Bounded optimal target weight: With In contrast to a great amount of works on chaotic systems, few results are combined the DHDP techniques into the design process. This paper focuses on a tracking control structure using DHDP as an online learning control element. Specifically, we design a control input (17), which contains the filtered tracking error and the cost function. Once the long-term cost function is minimized, an optimal control input will be generated. Now, the following theorem is explained to show the stability of the closed-loop system. Consider the discrete-time system given by equation (1). Let the control input is described as equation (17) and the adaptation law is shown by equation (16). Under the assumption 1, if the design parameters are chosen appropriately, the proposed algorithm can guarantee that all the signals are semiglobally uniformly ultimately bounded (SGUUB) and the tracking error Consider the following Lyapunov function
Assumption 1
Remark 1
Theorem 1
Proof
According to the filtered tracking error (equation (18)), the first difference of
Using the Cauchy-Schwarz inequality (equation (20)) and the equation
By subtracting
The first difference of
Substituting equation (22) into the above equation, we have
Based on the equation
Based on
Then, using the Cauchy-Schwarz inequality and the inequation
The first difference of
Based on the above analysis, let the first difference of the Lyapunov function be
Substituting equations (21), (27) and (28) into the above equation. By assumption, select
Finally, we get the first difference of
Given that conditions (equation (29)) hold and
6. Simulation example
In this section, an example is provided to demonstrate the effectiveness of the tracking control and the filtered tracking error is proposed in this paper. In order to illustrate the feasibility of the theoretic results, the Henon Mapping chaos model will be considered. In addition to a fuzzy controller will be designed based on the DHDP.
The control objective is to design an adaptive fuzzy controller for the system (equation (1)) that all the signals in the closed-loop system remain SGUUB, the system state vector
In this paper, the control system framework of the nonlinear tracking control with filtered tracking error has been implemented for DHDP design and analyzed using the Lyapunov stability approach, and a fuzzy logic system is applied to approximate the long-term cost function
The Henon Mapping chaos system has been described in equation (1), and the simulation parameters are set as λ1 = 0.25, τ
c
= 0.01, α = 0.5, c = 0.1 and k
v
= 0.5. The system initial states are chosen as
Based on the chaotic system (equation (1)) and the fuzzy controller (equation (17)), the simulation results are shown in Figures 2–4. The tracking performance is described in Figure 2 and it can be seen that a good tracking trajectory is achieved. The adaptive law of θ
c
is given in Figure 3. Last, Figure 4 shows the control input of the chaotic system and it can be observed from Figure 4 that the control input is bounded. It is easy to see that the proposed fuzzy controller make the closed-loop system stable.
x2(k) (solid line) and x
d
(k) (dashed line). The adaptive law θc. The control input u .


7. Conclusion
This paper makes full use of the DHDP design in the optimal tracking control setting with filtered tracking error. By using the reinforcement learning, an adaptive control was proposed to solve the tracking problem for the Henon Mapping chaotic system. The long-term cost function which is approximated by a fuzzy logic system is minimized so that an optimal control input can be achieved. Results have shown that the tracking error and the adaptation law were uniformly bounded in the sense of Lyapunov. A simulation example has been applied to verify the effectiveness of the method.
Footnotes
Funding
This research was supported by the Natural Science Foundation of China under grant number 61104017.
