Abstract
Designing observer-controller structures for nonlinear system with unknown dynamics such as robotic systems is among popular research fields in control engineering. The novelty of this paper is in presenting an observer-based model-free controller for robot manipulators using reinforcement learning (RL). The proposed controller calculates the desired motor voltages that fulfil a satisfactory tracking performance. Moreover, the uncertainties and nonlinearities in the observer model and RL controller are estimated and compensated for by using the Fourier series expansion. Simulation results and comparison with the previous related works (extended state observer and radial basis function neural networks) indicate the satisfactory performance of the proposed method.
Keywords
Introduction
Reviewing the problems associated with model-based controllers reveals the importance and necessity of model-free controllers. The first problem is identification of a suitable mathematical model for the system. Many systems are involved with some simple phenomenon such as friction. However, introducing a perfect dynamic model for friction is a challenging task. The second problem is sensing requirements. A model-based controller requires various feedbacks. Some of these signals are usually contaminated with noise such as acceleration feedbacks. Thus, filtering approaches are required which can increase the controller computational load and require further tunings. External disturbances, parametric uncertainty and un-modelled dynamics are other challenges which should be taken into consideration. To solve these problems, model-free control approaches have been presented (Karayiannidis et al., 2016; Lösch et al., 2018; Meng et al., 2018).
In recent years, various controllers such as adaptive fuzzy (Fateh and Khorashadizadeh, 2012; He and Dong, 2018), adaptive sliding mode (Lee et al., 2017), neural network (NN) (Izadbakhsh and Khorashadizadeh, 2020; Sun et al., 2016) and reinforcement learning (RL) (Li et al., 2017; Pfeiffer et al., 2017; Ren et al., 2020; Tai et al., 2017) have been used as model-free controllers for robotic systems. In the RL approach, the agents learn their policy and choose the control signal by trial-and-error exploration combined with reward signals coming from the environment (Falco et al., 2018; Lewis et al., 2012; Li et al., 2017). Recently, several important optimal control approaches have been presented along the RL algorithm for nonlinear continuous time systems. Using the RL approach to solve the online optimal control problem, a partially model-free RL algorithm was presented in the context of policy iteration for a linear continuous time system (Vrabie et al., 2009). Solving the linear quadratic tracking problem for continuous time systems has been the focus of many researches in recent years. According to Vrabie et al. (2009), an adaptive optimal NN controller was introduced for nonlinear constrained-input systems with unknown dynamics (Modares et al., 2013). To design the policy iteration technique, an NN actor–critic structure has been used. In Modares and Lewis (2014), an online integral RL-based algorithm has been developed. With respect to the aforementioned paper, many RL-based optimal control approaches were constructively introduced (Chen et al., 2019; Modares et al., 2015; Padmanabhan et al., 2019; Yang et al., 2019).
Combination of RL and adaptive fuzzy systems has been presented for a class of robot arms in Lin (2003). The action-generating element is a fuzzy approximator with a set of adaptive parameters which are calculated based on the adaptation laws obtained from stability analysis. In Tang et al. (2014), this action-generation algorithm is performed using NNs to control robot manipulators with unknown dynamics and dead-zone nonlinearities. It should be mentioned that the aforementioned RL-based approaches for robotic systems require all the system state variables such as position and velocity signals (Lin, 2003; Modares et al., 2013; Tang et al., 2014). Recently, many observer-based controllers have been presented in control engineering (Elkenawy et al., 2020; Wang and Deng, 2020; Wang and Pan, 2019; Wang et al., 2019). A cascade control structure based on a finite-time observer has been designed using the Lyapunov approach for tracking control of an autonomous under-actuated ship (Wang and Pan, 2019). The observer estimates the unmeasurable signals and the cascade controller has been designed with the aim of controlling the translation and rotation subsystems. A surge-heading guidance control system has been designed using a finite-time observer in the path tracking problem (Wang et al., 2019). However, many commercial robot manipulators provide just the position signals (Gholipour and Fateh, 2018; Nasiri et al., 2020). Thus, to solve this problem, an observer-based RL controller is presented in this paper.
Due to the excellent uncertainty approximation ability of function approximation techniques such as the Fourier series (FS) (Izadbakhsh et al, 2019), Szász–Mirakyan operator (Izadbakhsh et al, 2020), and Legendre polynomials (Zarei and Khorashadizadeh, 2019) much interest has been generated in this area (Chien and Huang, 2012; Kai and Huang, 2013). Similar to other estimators such as fuzzy systems and NNs, a function can be expanded by an FS. However, in comparison with neuro-fuzzy systems, the number of adjustable parameters in the FS expansion is less.
In this paper, an adaptive observer-based RL controller is proposed for robot manipulators using the FS expansion to construct the action generation process and critic structure. The superiority of this paper compared to Tang et al. (2014) is in introducing an observer-based actor–critic controller which eliminates the need for measuring all of the state variables. In addition, the control signal in the proposed method is motor voltages, while in Tang et al. (2014), the actuators have been excluded which may deteriorate the controller performance in high-speed applications (Huang and Chien, 2010; Khorashadizadeh and Sadeghijaleh, 2018). Also, to decrease the computational load of the controller, the FS expansion has been utilized instead of neuro-fuzzy systems. The adaptation laws for tuning the FS coefficients are obtained from the Lyapunov-based stability analysis. Moreover, in order to improve the observer performance, the nonlinearities originated from the manipulator dynamics are estimated and compensated for by using the FS expansion. It should be emphasized that in previous related works on RL, usually linear systems have been studied (Modares and Lewis, 2014; Vrabie et al., 2009), while in this paper, a multivariable nonlinear system is controlled using RL accompanied with stability analysis.
The remainder of the article is organized as follows: the second section presents a dynamic model of robot with direct current (DC) motors; in the third section, function approximation using the FS expansion is explained; the model-free observer and RL controller using the FS expansion are designed in the fourth section; stability analysis is presented in the fifth section; simulation results and comparisons are given in the sixth section; and the seventh section concludes the paper.
Modelling
Let us consider the motion Equation (1) of an electrically driven n-joint robot manipulator in the joint space (Spong et al., 2020):
where
in which
Substitution of
These matrices are indicated by boldface type to emphasize that they are related to the total robotic system. However, in order to make the design procedure and stability analysis simpler, the controller is developed based on the independent joint strategy (Spong et al., 2020). To be more precise, matrices in normal style correspond to a single joint or link. The state-space representation of Equation (5) for the
in which
FS expansion
It is well-known that FS expansion can approximate periodic functions. Using a limited time interval and assuming as a repeated function, a non-periodic function can be estimated by the FS. According to Khorashadizadeh and Fateh (2017), the non-periodic function
where
Considering the first
from which we obtain Equations (14) and (15):
where
Designing observer and RL controller
The block-diagram of the proposed system, including the proposed observer and the actor–critic controller is shown in Figure 1.

The structure of proposed method.
Observer
According to Equation (13), the uncertainty
where
Substitution of Equation (16) into Equation (9) yields Equation (17):
Let us consider
where
Now, define the following state observer as given by Equation (19):
in which
in which
In fact, Equation (20) describes the observer error dynamics. Define the error vector
in which
Therefore, to prove the stability and obtain the adaptation rule and observer, the strictly positive real (SPR)-Lyapunov technique has been used (Wang et al., 2004, 2010). Equation (23) can be rewritten as Equation (25):
in which
where
in which
RL controller
The proposed model-free controller using RL is formulated as follows. In fact, the actor and critic structures are based on Tang et al. (2014). However, the FS expansion has been utilized instead of NNs to make the tuning process simpler and reduce the computational burden of the controller.
Actor design
The desired closed-loop state-space Equation (29) is assumed as:
in which
where a positive constant vector is
Let us now define Equation (33)
as uncertainty of the controller for each link. Because
as given by Equations (34) and (35):
Similar to Equation (16),
Substitution of Equation (35) and (36) into Equation (32) results in Equation (37):
Using the definition
Also, using Equation (33), Equation (38) can be rewritten as Equation (39):
Substitution of Equation (34) into Equation (39) results in Equation (40):
Using the definition
in which
The control term
in which
Therefore, similar to the “Observer” subsection, to prove the stability and obtain the adaptation rule and controller, the SPR-Lyapunov technique has been used (Wang et al., 2004, 2010). Equation (44) can be rewritten as Equation (46):
in which
where
in which
Critic design
According to Tang et al. (2014), the reinforcement signal (long-term cost function) produced by the critic network is defined as given by Equation (50):
in which
Since
where
Stability analysis
In order to perform the stability analysis, two theorems are stated. The first theorem relates to the controller tracking error and closed-loop signals. The second theorem discusses the observer convergence analysis.
in which
In fact, it has been assumed that the actual and optimal values of weight vectors (free parameters) the actor (
in which
Substituting Equations (48), (51) and (52) into Equation (56) results in Equation (57):
Substituting Equations (50), (53) and (54) into Equation (57) yields Equation (58):
which can be simplified to Equation (59):
In other words, by combining Equations (58) and (59) we obtain Equation (60):
The term
It is obvious that the fourth and eighth terms in Equation (61), that is,
According to Assumption 1, it can be concluded that we obtain Equation (63):
The time-varying term
Due to the negative-ness of
In other words, this results in Equation (66):
Based on Remark 2,
By substitution of Equation (67) into Equation (66) and using
In other words, this results in Equation (69):
Thus, according to Equation (69), Equation (65) can be guaranteed and consequently, Equation (64) is satisfied and we have Equation (70):
Now, using Barbalat’s lemma, the asymptotic convergence of
Consider
From Equation (48), it follows that
Now, assume that
This completes the proof of Theorem 1.
in which
In fact, it has been assumed that the actual and optimal values of weight vector (free parameters) of the observer (
where
By substituting Equation (20) into Equation (74), Equation (74) can be rewritten as Equation (75):
Since
where
Appling Equation (72), Equation (77) is changed into Equation (78):
Now, it should be verified that we obtain Equation (79):
In order to guarantee Equation (79), one can guarantee the following inequality Equation (80):
According to Remark 1, to compensate the observer error, the term
As a result, Equation (80) can be rewritten as Equation (82):
In other words, we have Equation (83):
Based on Remark 1,
Thus, Equation (84) implies that Equation (79) is satisfied and consequently we obtain Equation (85):
Now, using Barbalat’s lemma, the asymptotic convergence of
Consider
From Equation (27), it follows that
Now, assume that
This completes the proof of Theorem 2.
Now, to guarantee the stability of the complete system with the proposed observer, consider the Lyapunov function candidate for each link as given by Equation (87):
in which
Simulation
The proposed observer-controller structure is applied on a two-link robot manipulator. Also, a comparison with the extended state observer (EOS) is presented to show the superiority of the proposed method.
The performance of the proposed method
The proposed controller Equation (36) and the observer Equation (19) are simulated using a 2-joint robot manipulator with permanent DC motors the symbolic representation of which is shown in Figure 2. The motor parameters and also matrices

The schematic of robot manipulator.
Using pole placement, the gain vectors
And for the second link are
The parameter
The coefficients of the FS in the actor, critic, and observer are initialized randomly in the range

The tracking performances of proposed method.

The tracking error of proposed method.

The control effort of proposed method.

The robot joint torques using the proposed method.

The observer error of proposed method.

The tracking error of proposed method in the presence of disturbances.

The control effort of proposed method in the presence of disturbances.
Comparison with EOS
In order to highlight the superiority of the proposed observer-controller with previous related works, the ESO (Talole et al., 2009; Zhang et al., 2020) is applied to the described robot manipulator. It is a model-free approach in which the lumped uncertainty is considered as an augmented state of system. Then, a linear state observer is used to estimate the states of system. Using Equations (9) to (11), and supposing the uncertainty
in which
According to Talole et al. (2009), the linear state observer is defined and the gain vector
Consider the following control signal Equation (92) (Talole et al., 2009):
where

The tracking performances of extended state observer.

The tracking error of extended state observer.

The control effort for extended state observer.

The observer error of extended state observer.
in which
where
Performance comparisons.
Comparison with radial basis function neural networks (RBFNs)
Here, in order to show the advantages of the proposed method in comparison with some previous works (Tang et al., 2014), instead of the FS expansion, RBFNs are utilized as actor and critic in the RL controller. RBFNs use the states of system in the regressor vector. In fact, the observer is eliminated in this simulation with the aim of focusing on the comparison of the FS expansion and RBFNs. Therefore, the main states of system have been used in both controllers to produce the control signal. Figure 14 shows the schematic of the RBFNs actor–critic controller. Using the RBFNs,

The structure of radial basis function neural networks reinforcement learning controller.
where
Therefore, Equation (96) can be rewritten as Equations (98) to (100):
in which
Two RBFNs are required as actor and critic parts in the controller. The stability analysis of this method (RBFNs) and obtaining the adaptive laws are similar to the proposed method. In order to perform the comparison, consider the following cost function Equation (101):
Simulating the same robot described in the “The performance of the proposed method” subsection, the aforementioned cost function for this controller, as shown in Table 1, takes the value of

The tracking performances of radial basis function neural networks-based method.

The tracking error of radial basis function neural networks-based method.

The control effort for radial basis function neural networks-based method.
Conclusion
An adaptive observer-based RL controller for an electrically driven robot manipulator has been presented in this paper. A model-free observer has been proposed in which the lumped uncertainty is estimated using the FS expansion. The controller is designed based on RL using an actor–critic structure. Furthermore, to generate the reinforcement signal in the critic and producing the control command in the actor, the FS has been utilized. The advantage of the FS in comparison with other possible choices such as neuro-fuzzy systems is fewer tuning parameters, and consequently less computational burden. The coefficients of the FS expansions are calculated based on the adaptation laws obtained in the stability analysis. Simulation results on a two-link robot manipulator actuated by permanent magnet DC motors show the efficiency of the proposed method in reducing the observer and controller tracking errors. In addition, comparisons with the extended state observer and RBFNs have been performed to show the superiority of the proposed method.
Footnotes
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
