Abstract
With the increasing connectivity among computational cyber-connected elements and physical entities, a unified representation that captures the interrelationship between the cyber and the physical systems becomes increasingly important. In this paper, we propose a novel representation for developing cybersecurity schemes for physical systems wherein the cyber system states affect the physical system and vice versa. Subsequently by using this representation, an optimal strategy via Q-learning is derived for the cyber defense in the presence of an attack. Since the cyber system under attack will affect the physical system stability and performance, an optimal controller by using Q-learning is considered for the physical system with uncertain dynamics. As an example, cyber-attacks that increase the network delay and packet losses are considered and the goal of the proposed cyber defense and optimal controller is to thwart the attack and mitigate the performance degradation of the physical system due to increased delays and packet losses. An illustrative example is given where the proposed theory is evaluated on the yaw-channel control of an unmanned aerial vehicle. Simulation results show that on the cyber side, both the attacker and the defender gains their greatest payoff whereas on the physical system side, the optimal controller is able to maintain the linear system in a stable manner when the cyber state vector meets a certain desired criterion.
1. Introduction
Cyber-physical systems (CPS) refer to engineered systems constructed as networked interactions of physical and computational cyber components. 1 Examples of CPS can be found in areas as diverse as automobiles, air transportation, civil infrastructure, power grid, embedded medical devices, and consumer appliances. Recently, with the development of information technology (IT) such as IT management and networking growth, the security in CPS has received attention. Moreover, as cyber and physical capabilities are becoming increasingly intertwined, a comprehensive framework that models the cyber system, the physical plant dynamics, and their interrelationship is also increasingly needed.
In general, there are two types of the representations for the security analysis of CPS in the existing literature: one that models the effect on the cyber systems under a certain specific attack;2–6 and the other includes the effect of cyber-attacks on physical systems.7–12 The former effort explores the behavior of the attacker as well as the defender, formulates the cyber changes under attacks, and presents appropriate strategies that bring the cyber system back to normal. For example, the study by Baumman and Sandmann introduces denial of service (DoS) flooding attacks by a continuous-time Markov chain and utilizes the state space method to compute security measures accurately. 2
Different from Baumman and Sandmann, 2 Zhu and Basar studied the cyber defense by modeling the actions of the attacker and the defender as a stochastic zero-sum game. 3 In the study by Ten et al., 4 the measure of vulnerabilities in cyber-physical systems with application to power systems is defined and a security framework including anomaly detection and mitigation strategies is provided. Sallhammar et al. evaluated cybersecurity by computing the expected probabilities of the attacker and using the probabilities to build a transition model through a game-theoretic approach. 5 In the investigation by Aenes et al., 6 the cyber vulnerability was dynamically evaluated by using a hidden Markov model which provides a mechanism for handling sensor data with different trustworthiness. However, this type of representation mainly focuses on the cyber system and neglects the fact that the states of the physical system also affect the cyber defense strategy.
In contrast, others concentrate on characterizing the dynamics of the physical system under attacks by extending the classic state-space description in order to include the attacks.7–12 For instance, in the report by Kwon et al., 7 the system dynamics include an extra term to model the deception attack. In the study by Liu et al., 8 the system state under attack is represented with an additive term, where the additive term is used to simulate the false data injection attack. Unlike Liu et al., 8 Teixeira et al. characterize the deception attackers by a set of objectives and propose policies to synthesize stealthy deceptions attacks in both linear and nonlinear estimators. 9 In the investigation by Fawzi et al., 10 the estimation and control of linear systems when sensors or actuators are corrupted by an attacker is provided, together with a secure local control loop that can improve the resilience of the system. On the other hand, Amin et al. define the control input under attacks as the product of the given input and a coefficient to characterize the effect induced by the DoS attacks. 11 A class of human adversaries, who are called correlated jammers, was considered by Zhu and Martínez. 12 By modeling the coupled decision making process as a two-level receding-horizon dynamic Stackelberg game, the authors propose a control law and analyze the performance and the closed-loop stability under attacks.
However, there are many weaknesses in the above reported works. 13 First, the representation can only describe a single type of attack due to the fact that attacks affect the system dynamics in a variety of ways. In particular, Pasqualetti et al. proposed a unified framework that is able to detect attacks; 13 however, it still has the two drawbacks mentioned next. Second, it is difficult to implement the representation developed in the literature so far since the system dynamics under attacks are considered known. For instance, due to random delays and packet losses caused by certain cyber-attacks, the physical system dynamics can be uncertain. Last but not the least, these representations fail to take the interactions between the cyber defense policy and the system controller under consideration.
In summary, to the best knowledge of the authors, little effort has been carried out in the literature to develop a representation that precisely characterizes the interplay between the cyber and the physical systems. Such a representation is necessary because inadequate decisions can be made for the cyber defense if the physical states are ignored. Likewise, the physical plant may not be stable if the controller is designed without considering the impact due to the changes in the cyber system.
Therefore, in this paper, we propose a framework for cyber-physical systems to (i) study optimal defense to mitigate attacks and (ii) to derive an associated optimal control policy for physical systems. First, we introduce a mathematical representation for the cyber-physical system, in which it was shown that the activities of the cyber system affect the states of the physical system and vice versa. Then based on this representation, we derive the optimal strategies for the defender and the attacker by considering them as two players in a zero-sum game. Since the cyber state influences the behavior of the physical system, next, an optimal controller for the physical system in the presence of uncertainties induced by the cyber system is revisited, based on that reported by Xu et al. 14 In addition, a condition on the cyber state vector is derived under which the physical system is stable. Finally, an illustrative example is given in which we show that on the cyber side, both the attacker and the defender gain their greatest payoff while on the physical side, the optimal controller is able to maintain the plant stable when the state vector of the cyber system meets a certain condition.
Thus, the main contributions of this work include: (i) a novel and comprehensive representation of the cyber-physical system that captures the interrelationship between the cyber and the physical elements; (ii) the development of the optimal strategies for the defender and the attacker; (iii) the application of the optimal controller for the physical system in the presence of uncertain dynamics induced by the cyber system; 14 and (iv) the demonstration of how the proposed theory can be applied to the control of the yaw-channel of an unmanned aerial vehicle (UAV) in the presence of an attack.
The rest of this paper is organized as follows. The proposed representation for the cyber-physical systems is introduced in Section 2. In Section 3, the optimal defense and attack policies are derived and presented, followed by the optimal controller design for the physical system introduced in Section 4. The illustrative example including policy derivation as well as the simulation results are presented in Section 5 and this paper is concluded in Section 6.
2. Proposed representationfor cyber-physical systems
In this section, the proposed framework for the cyber-physical systems is introduced. Figure 1 depicts the proposed representation for the optimal defense/control scheme.

Proposed representation of a cyber-physical system.
2.1. Cyber system
Consider the cyber system described by a nonlinear discrete-time system given by
where
The cyber state
In particular, we propose a more concrete representation of the cyber state as
where
As depicted in Figure 1, the cyber system output in the proposed representation is described as
where
One can assess the health condition or even the specific attack on the system by exploiting the cyber state
It is important to note that the physical system state
where
Threshold form:
Linear form:
Quadratic form:
In this paper, the attacks considered will increase the network delay and packet losses which in turn will make the linear time-invariant system as an uncertain stochastic time-varying system. The goal of the cyber defense and optimal controller is to mitigate the increase in random delays and packet losses and performance degradation of the physical system.
2.2. Physical system
As shown in the right block in Figure 1, the physical system is described as a linear discrete system in the presence of a disturbance given by
where
It is important to note that unlike the classical linear discrete system, the system matrices described by equation (4) are a function of the cyber state
In conclusion, the cyber state vector, whose update is subject to the attack/defense decisions, changes the physical system dynamics. As a result, the control input needs to be adjusted to drive the physical states back to the desired value. The changes in the cyber and physical states, in turn determine the cyber output and hence the attack/defense decisions. A summary of the interrelationship between the cyber and the physical systems is shown in Figure 2.

Interrelationship between the cyber and the physical system.
Hence, the objective is to design an optimal policy by using a cost function for the physical system with unknown system dynamics induced by the cyber system. Therefore, by (i) including the physical system state in the assessment of cyber health condition and (ii) considering the influence on the physical system dynamics induced by the cyber states when designing the optimal controller, the proposed optimal defense/control scheme offers a coupled design which is able to capture the influence of the cyber and the physical systems.
3. Optimal attack/defense policyfor cyber systems
In this section, the optimal attack and defense policies for the cyber system are derived, while in the next section we derive the optimal controller for the physical system with the presence of the delay and packet loss. We also derive the condition for the delay and packet loss under which the physical controller can be stabilized. The optimal controller gain will be computed and applied to the physical system once the delay and packet loss satisfy the condition. Otherwise appropriate defense strategy needs to be launched in order to drive the cyber states (delay and packet loss) to meet the criterion.
In this section, we first model the interactions between the attacker and the defender as a two-player zero-sum Markov game. 16 Then after defining the instant payoff as well as the expected discounted payoff function, we introduce two lemmas to show the existence of the solution of the game and the optimal policy. Next, the Q-function is proposed and it is shown in Theorem 1 that, using the Minimax-Q algorithm, 17 the Q-function converges to the game value. As a result, the optimal strategies for the defender and the attacker in order to gain their greatest discounted payoff are also derived.
Consider the cyber system with dynamics described by equation (2) and an output function in quadratic form of the state vectors, i.e. as
where the cyber state vector
Let
As illustrated in Figure 3,

Each subset of
Let
Specifically, we let the instant reward be defined as
which consists of the cost of the cyber state, physical state, defense, and attack. The defense cost is defined as
After introducing the definition of the instant payoff, we now consider the expected discounted payoff function over multiple stages. Let
where
where
Based on these two lemmas, we use an iterative Q-learning method to search for the game value
Accordingly, the optimal action dependent value function
From equations (9) to (11), one can conclude that if the action pair sequence
The Minimax-Q algorithm proposed by Littman is adopted to obtain
where
where
The proof of Theorem 1 is similar to the theorem given by Littman. 21
In addition, since
A flowchart of the proposed method to obtain the optimal defense/attack strategy is shown in Figure 4.

Flowchart of the optimal policy for the defender/attacker.
4. Optimal controller design
In this section, we introduce the optimal control scheme for the physical system based on the previous work. 14 First, we model the linear discrete-time system with dynamics that is unknown and altered by the cyber state vector, which includes packet losses and time delays since these are two important metrics for the network that may cause deterioration or potential instability of the system. 22 We then introduce the optimal control gain and show that the system is stable only when the cyber state vector satisfies a certain criterion. The cyber system needs to launch the appropriate defense if its state vector fails to satisfy the criterion. The development of the system dynamics as well as the Q-function update law is taken from the paper by Xu et al. 14 In summary, we show that the cyber state vector affects the optimal controller design and meanwhile the states of physical system also have an impact on designing the defense for the cyber system.
In cyber-physical systems, there are two types of network-induced delays: the sensor-to-controller delay and the controller-to-sensor delay. With the assumption that the former is negligible, the linear continuous system can be described as 14
where
where
where the system matrices are a function of the unknown random delays, and packet losses or the cyber state vector which are given by 14
and
where
Consequently, the optimal control gain is represented in terms of
Now define the residual or temporal difference error as
Next, we define an auxiliary residual error vector as
Similarly, the dynamics of the auxiliary vector are derived as:
It was shown by Littman that,17 with the update law of equation (22), there exists a positive constant
Finally, we show the sufficient condition on the cyber state in term of the delay and packet loss that need to satisfy in order to maintain the system to be stochastically stable. Consider the systems with slowly-varying parameters, since the initial stabilizing control and disturbance inputs are given, the linear discrete-time system can be represented as
According to the definition of stability for stochastic linear time-varying system,
24
if eigenvalues of
Then
where
Combining equation (23) with equation (24), we have
Therefore, in order to maintain stability, the expected values of the delays and packet losses should satisfy
where
5. An illustrative example
In this illustrative example, the proposed framework is verified on a small-scale UAV helicopter with remote controller. The objective of the controller design is to stabilize the yaw rotation rate with the presence of two types of cyber-attacks. The attacker aims to maximize the payoff, which are given in terms of the network delay and packet losses in this case, such that the yaw channel becomes unstable. The defender, on the other hand, aims to limit the delay and packet losses under a certain threshold. We will show that on the cyber side, both the attacker and the defender gain their greatest payoff while on the physical system side, the optimal controller is able to maintain the yaw rate stable when the cyber state vector expressed as delay and packet loss meets the derived condition.
5.1. Physical system setup
In this illustrative example, we consider the control of the yaw rotation of a small-scale UAV helicopter. A yaw rotation, as illustrated in Figure 5, is a movement around the yaw axis of a rigid body that changes the direction it is pointing. 26 The yaw rotation control is one of the most challenging tasks in controlling small-scale UAVs because even a small control input or disturbance can cause the vibration of the light-weight body. 26 Since it has been verified that the yaw-channel dynamics for small-scale helicopters can be physically decoupled from other channels,27,28 it is reasonable to assume that the yaw-channel dynamic is a single-input–single-output system. Furthermore, after applying the prediction-error method, 29 an accurate fourth-order model is proposed as 26
where

Illustration of a yaw rotation.
The other parameters of the physical system are introduced as follows. The total simulation time is 200 steps with the sampling time of 100 ms and the positive constant
5.2. Cyber system setup
As illustrated in Figure 6, we suppose that the UAV is controlled by a base station through a wireless network that suffers from cyber-attacks. As stated earlier, we choose packet losses

Diagram of the UAV with remote controller.
Smurf attack is an example of amplification distributed denial of service (DDoS) attack that exploits the unprotected networks to generate significant traffic load on the victim network.30,31 Slow read attack, on the other hand, tries to exhaust the server’s connection pool by sends legitimate application layer request but reads the response slowly.
32
Based on these characteristics, we model the delay and packet loss rate to increase exponentially under the smurf attack and linearly under the slow read attack, which are illustrated in Figure 7(a) and (b). Furthermore, the corresponding strategies that are capable of defending smurf attack and slow read attack are denoted as

Models of delay/packet loss rate under (a) smurf attack, no defense; (b) slow read attack, no defense; (c) smurf attack with the corresponding defense; (d) slow read attack with the corresponding defense.
The cyber output is defined as
Next, as presented in the flowchart in Figure 4, we divide the cyber output
It is important to note that we make
The system information for this particular example is summarized as in Table 1. The simulation is performed with the algorithm described in Figure 4 and numerical values shown in Table 2.
Summary of system information used in the illustrative example.
Numerical values used in the simulation.
5.3. Simulation results
In the simulation, the optimal defense/attack policies for the cyber system and the optimal controller are derived in the presence of delay and packet losses. Since the delay and packet losses are generated from the cyber system, they are determined directly by the policy launched by the defender. After deriving the optimal defense/attack policies, two scenarios are considered in the simulation. In the first scenario, we let the defender launch the cyber defense policy based on the probability distribution given by the derived optimal policy. By contrast, in the second scenario, the defender selects the defense actions at random.
5.3.1. Results of deriving the optimal attack/defense policies
First, we shall show the simulation results of deriving the optimal attack/defense. After about 2000 iterations, the Q-values for all action pairs converge to fixed values. To avoid redundancy, we only show the Q-values for the attacker and the defender in region

Q-values in region
Percentages for each action in the region.
It can be concluded from Table 3 that when
The proposed model and analysis is verified through the following simulation. We start the system with the cyber state initialized to zero and stop after 1000 iterations. During iteration, the attacker and defender will (i) determine which region

Evolution of the states (a): delay; (b): packet loss rate.
From Figure 9 it can be concluded that after a rapid increase at the beginning, the delay and the packet loss rate remains relatively stable so that the attacker gains the largest expected payoff in terms of the delay and packet losses. This is achieved by loading much more

Evolution of the output.

Evolution of average payoff.
In addition, the simulation is repeated for the case where the two attacks/defenses can be loaded simultaneously. As a result, a table similar to Table 3 is obtained except that two extra columns are added, which are the probability distributions of simultaneously loading two attacks

Evolution of the output, when two attacks/defenses can be loaded simultaneously.
5.3.2. Scenario I: defender chooses the optimal policy
In this scenario, we let the defender launch the defense policy based on the probability distribution given by the derived optimal policy. As a result, the delay and packet losses have been limited to relatively low values so that the system always stays out of the failed region, which is as verified in Figure 9(a). Consequently, equation (25) is satisfied in this scenario. The simulation results of the regulation errors for the physical system are shown in Figure 13, where the state regulation errors converge to zero thus forcing the closed-loop system being stable. Therefore, we show that on the cyber side, both the attacker and the defender gains their greatest payoff while on the physical side, the optimal controller is able to maintain the plant stable when the cyber state vector meets the derived criterion.

Regulation errors in Scenario I where the cyber defense is optimal.
5.3.3. Scenario II: defender chooses a random policy
In the second scenario, the cyber defense is selected at random rather than based on the optimal probability distribution given in Table 3. As a result, the attacker manages to compromise the system in some cases and the cyber states go far beyond the limit, as verified in Figure 14 in which the time delay is plotted. Consequently, equation (25) cannot be satisfied and thus the system becomes unstable. The regulation errors in this scenario are plotted in Figure 15, where it can be seen that the errors do not converge.

Delay in Scenario II where the cyber defense is randomly selected.

Regulation errors in Scenario II where the cyber defense is randomly selected.
In summary, the simulation results verify that that the decisions made on the cyber system have an effect on the convergence of the physical system. The system is stable when applying the optimal control in the physical plant and optimal defense policy in the cyber system. If the states go abnormal such that equation (25) is not satisfied, appropriate actions needs to be launched on the cyber system to bring them back to normal or the physical plant has to be shut down to avoid further damages.
6. Conclusions and future work
With the increasing meshing among the cyber-connected elements with the physical entities, the representation for such cyber-physical system becomes more complicated. In this paper, we have proposed a representation that captures the interrelationship between the cyber and physical systems such that the states in the physical system affect the decision made on the cyber systems and vice versa. Based on this representation, the optimal defense and attacks are given to gain the greatest payoff. An optimal controller from the literature is revisited to maintain the stability of the physical system in the presence of the uncertainties induced by the cyber state vector. Since the proposed representation is in a general form, it can be used in a variety of applications including autonomous systems. In particular, the cyber defender is able to make thorough decisions by selecting appropriate cyber state vector and output and customizing the payoff function that is of interest. Meanwhile, there are some recent works focusing on modelling and controlling for multi-agent networks or cyber-physical systems.33–35 For example, the work by Xue et al. characterizes a binary notion of security and characterizes security levels in terms of the graph matrix and its spectrum, 33 which is complementary to control-theoretic modeling of attacks in cyber-networks and networked control systems. Based on these works, as future work, we can consider studying the impact of different attacks on the network performance to generate a more accurate model for the cyber system dynamics.
Footnotes
Funding
This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.
