Abstract
This article proposes a novel dynamic response reconstruction approach for structural health monitoring using densely connected convolutional networks. Skip connection and dense block techniques are carefully applied in the designed network architecture, which greatly facilitates the information flow, and increases the training efficiency and accuracy of feature extraction and propagation with fewer parameters in the network. Sub-pixel shuffling and dropout techniques are used in the designed network and applied to reduce the computational demand and improve training efficiency. The network is trained in a supervised manner, where the input and output are the measurements of the available channels at response available locations and desired channels at response unavailable locations. The proposed densely connected convolutional networks automatically extract the high-level features of the input data and construct the complicated nonlinear relationship between the responses of available and desired locations. Experimental studies are conducted using the measured acceleration responses from Guangzhou New Television Tower to investigate the effects of the locations of available responses, the numbers of available and unavailable channels, and measurement noise. The results demonstrate that the proposed approach can accurately reconstruct the responses in both time and frequency domains with strong noise immunity. The reconstructed response is further used for modal identification to demonstrate the usability and accuracy of the reconstructed responses. The applicability of the proposed approach for structural health monitoring is further proved by the highly consistent modal parameters identified from the reconstructed and true responses.
Keywords
Introduction
Ensuring the safety and reliability of civil engineering structures is an important agenda in society. To real-time monitor and assess the structural condition and early report the anomaly in structural conditions and vibration behaviour, an increasing number of long-term structural health monitoring (SHM) systems have been installed on large-scale civil engineering structures, such as bridges and high-rise slender structures.1–4 However, during the long-term monitoring, it is possible that measurement data from one or more sensors could be lost due to the technical issues such as the malfunction of cable connections, power supply interruption, signal transmission disturbance, sensor malfunction or regular maintenances such as equipment check and sensor replacement. In practice, the locations of installed sensors on the monitored structures are selected carefully, and the number of sensors is much smaller compared to the total number of degrees of freedom (DOFs) of structures because of the limitations on the budget and available channels of data acquisition systems 5 and the inaccessibility of some locations for measurement during operations. 6 Credible damage detection and condition assessment techniques assume that all the sensors are functioning in perfect condition. 7 The incomplete acquired measurements that miss signals of one or more channels will greatly affect the performance of structural condition monitoring since some important local information of structures could be lost. For example, under the extreme events, SHM data could be lost due to the malfunction of systems and the lack of power supply of some sensors and data acquisition systems. However, the measurement data under those events are important for assessing structural conditions and evaluating the damage severities. As a result, structural dynamic response reconstruction of the lost data for some channels becomes an important research topic in the field of SHM.
In the past decades, several types of methods regarding the reconstruction of dynamic responses have been proposed utilizing the finite element model (FEM)-based techniques. Transmissibility concept–based 8 structural response reconstruction methods were studied in both the frequency domain 5 and the wavelet domain. 9 The responses at the desired locations are reconstructed using the responses at other measured locations using the transmissibility matrix. The limitation of these methods is that the information on the locations of the input excitations is necessary, and an accurate FEM is required. Deriving the transformation matrix utilizing the FEM of the structure to obtain the response at unavailable DOFs through the available response is another typical method. 6 This research is further combined with empirical mode decomposition (EMD), which decomposes the measurable time domain data into several intrinsic mode functions (IMFs) and reconstructs the responses at unavailable DOFs using the independent transfer equation of each IMF.10,11 The selection of IMFs is manually conducted, leading to the performance of the methods highly dependent on empirical experiences. An alternative structural response reconstruction method was developed based on optimal multi-type sensors placement and Kalman filter with unknown excitations.12,13 Minimum-variance unbiased estimates of the generalized state of the structure and the external excitations were computed based on the measurements of the carefully selected locations. The generalized state of the structure and the external excitations were then used to reconstruct the response at the desired DOFs that are not installed with sensors.
These FEM-based methods demand an accurate FEM of the target structure to calculate the transmissibility matrix, identify the mode shapes or determine the optimal sensor locations. However, developing an accurate FEM for large and complex civil engineering structures with a large number of elements and DOFs is time-consuming and challenging. Furthermore, system and geometry properties, and boundary conditions of the structures are changing with the variations of the environmental and operational conditions, which leads to the establishment of an accurate FEM representing the in-service structures much more difficult. However, with the rapid development and application of data analysis techniques, realizing the response reconstruction by the data-driven methods is attracting significant attention. The majority of these data-based methods focus on reconstructing the strain data14,15 or the strain distribution 16 using the inter-channel correlation. Compared to the strain data, acceleration response is more sensitive to the structural vibration and accumulated damage in structures. However, the nonlinear inter-channel relationships are more complicated. Studies on the reconstruction of vibration acceleration data are mainly conducted to reconstruct the randomly missing data or short-term continuous lost data in a specific channel. For example, Bao et al. 17 recovered the randomly lost acceleration data by leveraging the compressive sensing and further improved the developed method with group spare optimization 18 and machine learning techniques. 19 Yang and Nagarajaiah 20 also addressed the random data missing issue with sparse representation for inter-channel data reconstruction of low-rank structures. Wan and Ni 21 proposed a Bayesian multi-task learning method with multi-dimensional Gaussian process to predict the short-term continuous lost data using the measurements from the same sensor ahead of the forecasting. However, when a long-term successive acceleration data loss occurs or acceleration responses at critical but inaccessible locations are desirable, these methods may not be capable of accurately reconstructing the required responses.
With the advance of big data analytics in computer science and the tremendous growth of computing power, deep learning has gained much attention in broad research areas. Deep learning is capable of automatically identifying the pattern and mining the hidden highly abstract features from a large volume of data. SHM systems continuously measure the structural vibration responses, which can be used as training data and serve as a platform for using deep learning techniques to extract the abstract features and train the complex nonlinear relationships between input and output for dynamic response reconstruction. Developing and applying the deep learning techniques for SHM have gained significant research attention recently. The existing development and application of deep learning techniques in SHM mainly focus on the visual inspection and damage detection of structures using the crack images,22–24 vibration responses25–27 and pre-extracted vibration characteristics28,29 as the input to the networks. A comprehensive review of deep learning–based SHM techniques is provided by Ye et al. 30 Through training the networks with a massive amount of datasets, the high-level representations of the input data are automatically extracted by the deep learning models and then used to construct the nonlinear relationships with the output, which can be the structural crack locations and widths, damage location and severity. More applications in computer vision have been reported, that is, the pixel-to-pixel recovery of images and audio super resolution31,32 and image inpainting.33–35 Considering that the concept of structural response reconstruction is similar to one-dimensional image recovery, deep learning technique has high potential to be leveraged to conduct the response reconstruction for SHM. Fan et al. 36 developed the vibration lost data recovery approach by using convolutional neural networks (CNNs) for recovering the lost sensor data using the measurement data from the same sensor set under various loading conditions. However, the reconstruction of vibration responses from instrumented sensor locations to other locations with sensor data lost or without sensors, using the artificial intelligence techniques, has rarely been studied.
This article proposes a novel structural response reconstruction approach based on densely connected convolutional networks (DenseNets) for inter-channel vibration response reconstruction. Skip connection, dense connection, sub-pixel shuffling and dropout techniques are described and applied in the designed DenseNets. The effectiveness and accuracy of the proposed approach for the reconstruction of the vibration acceleration responses are investigated with the in-field testing data on Guangzhou New Television Tower (GNTT). The effects of the locations of available sensors and malfunction sensors, the number of available sensors and measurement noise are studied. The response reconstruction accuracy is evaluated in both time and frequency domains. Modal identification using the reconstructed responses is also conducted to validate the accuracy of reconstructed responses qualitatively and quantitatively.
Methodology
In this study, the designed deep learning model, namely, DenseNets, for structural dynamic response reconstruction is a one-dimensional full CNN with densely connected layers. The DenseNets learn the inter-sensor relationships by a supervised manner using the training datasets generated from measured vibration data before the sensors at desired locations become unavailable (faulty or inaccessible). The trained model can then be used to reconstruct the acceleration responses at desired locations leveraging the measurements from other available sensors. The idea of the architecture of this designed network is inspired from U-net 37 (also known as convolutional encoder–decoder 35 ), which consists of the compression path to capture the high-level features of the input and the reconstruction path to gradually expand the features to reconstruct the output. During the compression path, the feature maps are shrunken to facilitate the network to learn robust features of the input data while sifting the noise with uncertain patterns. Deep network with such architecture is proved to be efficient for conducting the pixel-to-pixel reconstruction tasks. 37 In addition, skip connection, dense connection, sub-pixel shuffling and dropout techniques are carefully applied to improve the performance of the developed DenseNets. Compared with the typical U-net, this newly designed network uses dense connection and skip connection instead of the successively connected convolutional layers to boost the information flow among the networks. The use of dense connection and skip connection strongly increases the training efficiency and accuracy of feature extraction and propagation, substantially reduces the number of parameters and alleviates the vanishing-gradient problem. In this case, the demanded volume of training data is significantly reduced, which makes the preparation of training data simpler. Fewer training data also bring efficient training compared with traditional deep neural networks. Moreover, optimized information flow also fastens the convergence speed which reduces the computational demand and further reduces the training time. Sub-pixel shuffling is a recently attractive method to replace the deconvolution which realizes the feature upscaling in a computationally efficient manner. Dropout technique is applied in the convolutional layers to mitigate the overfitting. The details and benefits of these techniques are elaborated in the following sections.
Configurations of the proposed DenseNets
The elaborated architecture of the proposed DenseNets is shown in Figure 1 and the detailed configurations of the proposed network are summarized in Table 1. Figure 1 shows an example of the networks using some functional SHM channel responses to reconstruct the response of a specific channel that is faulty or inaccessible.

The architecture of the proposed DenseNets.
The detailed configurations of the proposed DenseNets.
Leaky ReLU: leaky rectified linear unit.
The input and output of the DenseNets are the acceleration responses from n available locations and an unavailable location. The kernel number and size of each layer are given in Table 1. The first low-level feature extraction layer (Low level as shown in Figure 1) aims to enrich the shallow features from the original input data with limited available channels by applying the convolution operation. ‘Same’ padding in Table 1 means zero-padding equally at the beginning and end of the input to allow the layers generating output with the same length as input when the stride is 1 and with a half-length of the input when the stride is 2. From the shallow features, the higher representative features of the input are gradually extracted using three successive dense blocks (Dense 1–3 as shown in Figure 1). After each dense block, the length of features is halved to facilitate the network for refining the higher level representative features and eliminating the noise with no certain features. Meanwhile, the number of feature maps is doubled to store more information. The output of the last dense block (Dense 3) has the highest level abstraction of input and the largest number of feature maps as denoted in Figure 1. The reconstruction layers (R1 and R2 as shown in Figure 1) follow the last dense block to expand the feature maps using the sub-pixel shuffling. The sub-pixel shuffling used here doubles the length and halves the number of feature maps. After two reconstruction layers, the output feature maps have half of the length of the input data. The final convolutional layer (Final as shown in Figure 1) finally reconstructs the structural response data at desired unavailable locations by convolutional and sub-pixel shuffling operations without the nonlinear transfer of the output. 38 It should be noted that the output features at the symmetric levels have the same length, and the output features of the dense block are transported to the symmetric location by the skip connection. The shuttled features are concatenated with the output features of the reconstruction layers as the input of the next layer. It should be noted that the input and output sizes in Figure 1 and Table 1 are obtained, assuming that the length of data points in input and output sensor responses is 1024 time steps. Since the input data are halved three times by the dense blocks, the input length should be a multiple of 8 (23) to guarantee the length of the smallest feature maps is an integer. It is recommended to tailor the input response longer than the kernel size of the low-level feature extraction layer to fully utilize its capacity.
The proposed DenseNets for dynamic response reconstruction are based on supervised learning. For performing the structural response reconstruction at unavailable locations, the training datasets are generated using the measurement data with the functional sensors placed on these available and unavailable locations. The datasets consist of paired input data from available sensor locations and output data from unavailable locations. When using the training datasets with a massive number of samples, inputting all the samples to the networks at the same time will occupy a huge amount of memory and cause a large computational demand. Alternatively, the samples are loaded in batches to tune the parameters, that is, weights and biases of the proposed network. The network parameters are tuned using the back-propagation algorithm. During the training process, the input data flow through the DenseNets to reconstruct the responses, and the target of the training process is to minimize the difference between the predicted responses and true measured responses as output labels. The way to compute the loss between the predictions and labels is using the mean of the normalized L2-norm error of each sample in a batch as
where N is the number of samples in each batch, and
One-dimensional convolution
Compared to the traditional deep artificial neural network (ANN) that fully connects the neurons between adjacent layers, CNN is a class of deep neural networks consisting of one or more convolutional layers. CNN is a powerful deep learning architecture that can extract the hierarchical pattern in data by stacked convolutional layers with fewer parameters. A convolutional layer contains several convolutional kernels with each consisting of a weight matrix and a bias number to extract the feature of input. Through sliding the kernel across the input data with a specific stride, convolution results are obtained as the dot product of the weight matrix with the scanned segment of the input data plus the bias number. After the kernel reached the end of input, the results of all positions are concatenated together with the sequence to form the output feature map of this kernel. The output feature maps from all kernels will be subsequently used as the input for the successive layer. Generally, as the complex relationship between the input and output cannot be simply represented by linear functions, the output of convolution will be nonlinearly transferred by the activation function which can be expressed as
where O and I are the output and input of the convolutional layer, respectively; W and b are the learnable weights and bias of the kernel, respectively; and H and
Leaky ReLU has a broader range of output which mitigates the gradient vanishing. Meanwhile, it allows a non-zero small gradient to overcome the defect of ReLU, that is, the neuron is potentially never activated once a large gradient passes through it. 39 Acceleration data are considered as time-dependent one-dimensional signals. Therefore, the one-dimensional convolution is selected among all the layers in this application for structural response reconstruction. One-dimensional convolution can be considered as a special case of the two-dimensional convolution, where the convolutional kernel slides along only one axis of the input. To adopt this method for the case with multi-channel signals, the input acceleration data from multiple channels are parallelly arranged. Meanwhile, the kernel height is shaped as the same as the number of channels.
Skip connection
It should be mentioned that when tuning the networks with deep architectures of a layer-to-layer sequential connection using the gradient and back-propagation, the gradient vanishing in the deeper layers is a crucial problem, which significantly influences the extraction of higher level hidden features. 41 Skip connection directly shuttles the low-level features extracted in the bottom layers to the top layers and allows the features to be back-propagated to the bottom layers directly, which enhances the information flow and alleviates the gradient vanishing issue. 42 Moreover, it can also pass the low-level details that are lost during the convolution to the bottom layers. 35 Networks with skip connections are also known as residual neural networks (ResNet). The skip connection technique in the proposed DenseNets is used in both shuttling the features from the bottom layer to the top layer and embedding in the dense block. The involvement of skip connection significantly strengthens the feature propagation and consequently boosts the convergence efficiency.
Dense block
The dense block is used in the proposed network to enhance the extraction of the hidden high-level features from the input. A dense block consists of multiple densely connected convolutional layers.41,43 The dense connectivity strategy is implemented by skip connection between any layer and all subsequent layers as shown schematically in Figure 2. Dense connection improves the information flow from the shallow layer to the deep layers in this block and reduces the number of parameters by sharing the features of shallow layers with deep layers. The dense connection extracts the higher level features much more efficiently with fewer parameters than the traditional layer-to-layer sequential connected convolutional layers which use a large number of kernels and parameters. Compared with traditional deep learning models using the layer-to-layer connection, networks equipped with densely connected layers are stronger in feature extraction. Networks with dense connections are more efficient in training, and the extracted features are complete and comprehensive. Traditional layer-to-layer connected CNN with feature loss is hard to achieve the same training accuracy as the proposed networks equipped with the skip and dense connections, especially for time domain response reconstruction that belongs to pixel-to-pixel tasks which contain a vast feature extraction and propagation work, which is almost impossible using the traditional CNN. Layer-to-layer connection for response reconstruction requires much more kernels, resulting in a significant increase in the number of parameters to be trained. In this study, to achieve the similar results as the developed DenseNet, the required number of parameters for networks with the same depth but composed of layer-to-layer connection only is around two times and the training time is more than four times because it needs more training iterations for learning response features.

A schematic four-layer dense block.
The detailed configuration of these three successive convolutional layers is shown in Table 2. In each of the dense blocks, the number of kernel and kernel size are kept the same. The first three convolutional layers (Conv 1–3) implement the convolution by sliding the kernels with a stride of 1 for the output to have the same length. The last convolution layer (Conv 4) conducts the convolution utilizing the extracted feature maps from all the previous layers. Convolutional kernels slide with a stride of 2, which halves the dimension of feature maps. The high-level features of the output are used as the input of the next dense block.
The detailed configurations of the dense blocks.
Leaky ReLU: leaky rectified linear unit.
Sub-pixel shuffling
To gradually expand the feature maps in the reconstruction layers of the designed networks, the sub-pixel shuffling operation is embedded in the reconstruction layers following the nonlinear activation operation to upscale the feature maps by a factor of 2. Figure 3 demonstrates an example of sub-pixel shuffling with four feature maps consisting of six features each. It divides the feature maps to two groups and combines two feature maps in the same order of each group as one by interpolating one to another. It is a one-dimensional case of the sub-pixel convolution layer. 38 This is an efficient operation for upscaling features, which costs less computational demand than the general deconvolution. It has also been attested to have strong workability that introduces fewer artefacts in the output. 44

The schematic procedure of one-dimensional sub-pixel shuffling with four feature maps and six elements in each map.
Dropout
Overfitting is a common issue in the training of networks with deep architectures. It happens when the network model overfits the training datasets and loses the generalization. Instead of spending vast time and computational resources on training several networks with random initialization and averaging their outputs, dropout technique is an emerging tool for addressing the overfitting issue. Dropout technique de-activates a certain proportion of neurons and disconnects these neurons with the adjacent input and output layers to break up the co-adapted sets of neurons. 45 The selection of dropping out neurons is random when training with a different batch of samples. Those remained neurons are trained more robustly, and the generalization capacity of the networks is enhanced. Each neuron has an independent probability p to be dropped, where p = 0.5 is suggested by the inventors 45 and used in the training of the proposed DenseNets.
Experimental studies
GNTT and its SHM system
GNTT is a 610-m supertall slender structure. It has a tube-in-tube structural design including a reinforced concrete inner tube and a steel outer tube, which consists of 24 concrete-filled tube columns and 46 steel ring beams and bracings. Twenty-four columns are uniformly spaced in an oval shape and twisted in the vertical direction as shown in Figure 4(a). The size of the oval changes from 50 × 80 m2 at the ground level to 20.65 × 27.5 m2 at the height of 280 m and further to 41 × 55 m2 at the top of the tube with a height of 454 m. Thirty-seven floors connecting the inner tube and the outer tube by beams are used for offices, entertainment, catering and emission of the television signals. 4 A sophisticatedly designed SHM system including more than 600 sensors is installed on the GNTT for monitoring the structural vibration behaviour and ambient environmental conditions in both the construction and in-service stages. The in-service acceleration measurement data are selected as the training and testing of the proposed method. The location placement of the accelerometers is shown in Figure 4(b). Totally, 20 uni-axial accelerometers are installed at the assigned heights in both the long-axis and short-axis directions. The sampling rate of the acceleration measurement is 50 Hz. The raw measurement is processed by a high-pass filter with a 0.05-Hz cut-off frequency to eliminate the shift at the zero frequency. The previous studies46,47 have validated the effectiveness of the SHM system and extended GNTT as a benchmark platform for high-rise structures. The vibration measurement data will be used to validate the effectiveness and accuracy of the proposed DenseNets for dynamic structural response reconstruction.

(a) GNTT and (b) the deployment of accelerometers and sensor numbers.
In this study, 24-h measurement data along the short-axis direction including channels 1, 3, 5, 7, 11, 13, 15 and 17 are processed as the datasets for evaluating the accuracy of using the proposed approach for response reconstruction. Two case studies are conducted. The first case study is conducted to investigate the effectiveness of reconstructing the response of one unavailable channel using the responses of different numbers of available channels. The second case study is conducted to evaluate the performance of the proposed method, when data from more than one channel are lost and are reconstructed, especially when the most correlated channel to the desired channel is not available. The noise immunity of the proposed method will also be investigated in the second case study using measurement data contaminated with noise effect.
Data pre-processing
The previous studies47,48 have conducted the modal identification of the GNTT and reported that the first 15 vibration modes are within 2 Hz. Therefore, the frequency of interest of this study is selected as 0–2 Hz. The original measurements sampled at 50 Hz are pre-processed using the low-pass filtering with a cut-off frequency of 5 Hz and then downsampled to 10 Hz. The filtering processing effectively eliminates the redundant information contained in the signals and consequently reduces the number of features to be learned by the network. Meanwhile, downsampling the acceleration response by a factor of 5 significantly reduces the dataset size and greatly boosts the training efficiency. The pre-processed acceleration data are then scaled to a range between −1 and 1 without affecting its statistical distribution by
where A is the acceleration response vector after downsampling, which consists of the measurements from totally k involved channels with n sampling points;
The normalized data A′ are then used to generate the training, validation and testing datasets. For training the proposed DenseNets by supervised training, the input and output are the measurements from available sensor locations and the desired unavailable locations. To facilitate the network to learn robust features from comprehensive datasets, the normalized acceleration responses in 24 h are divided as 96 segmentations, where each segmentation contains 15-min measurement data. The last 10% of segmentations are used as the testing data to simulate the situation that the measurement data of the desired channels are unavailable. This will be used to test the accuracy of the network for dynamic response reconstruction. The remaining segmentations are randomly grouped as the training and validation datasets containing 80% and 10% of the total number of segmentations, respectively, to construct and validate the relationship between input and output. As mentioned in section ‘Configurations of the proposed DenseNets’, utilizing all the training data to tune the parameters of deep networks at the same time is impractical. Due to the change in load and vibration characteristics with the variation of ambient conditions, inputting the segmentations one by one will lead to the model being unstable and hard to reach the optimal parameters. Alternatively, the segmentations are cut as small samples with 1024 data points each, by a sliding window which is 1024 data points long with a stride of 512. Those samples are shuffled and inputted to the DenseNets in batches for both training and validation. The batch size is chosen as 32 for the following studies. The acceleration data with measurements from available channels are directly inputted to the trained networks, and the measurements of the desired channel are defined as the labelled output. For testing the network, the reconstructed responses are compared with the original true measurement data from the channels at unavailable locations to validate the accuracy of the proposed approach.
The effect of input channels
Selecting proper available channels as the input of the DenseNet is important for an effective response reconstruction. In the first case study, channel 17 is assumed to be faulty after a period of service for making reliable measurements. The effect of using different locations and numbers of available channels as the input of the DenseNets to conduct the dynamic response reconstruction is investigated. First, the selection of the most correlated input channels for reconstructing channel 17 is studied. For response reconstruction by deep learning models, a channel containing substantial overlapped natural frequencies and vibration modes with the desired channel will have a higher correlation. Traditional calculation of correlation function of two response series in time domain may not be directly useful since the similarity in learned features, that is, vibration modal information between the recorded and desired channels, is more important for response reconstruction. However, manual selection of correlated channels may introduce large errors due to noise and measurement errors. Therefore, the correlations between channels are evaluated by one-to-one channel reconstruction. Recorded channels used for response reconstruction producing lower reconstruction errors means the responses from these channels can provide more useful information for feature learning and dynamic response reconstruction of desired channels. Seven groups of training, validation and testing datasets containing the available measurement data from one of the channels 1, 3, 5, 7, 11, 13 and 15, respectively, and the output from channel 17 are generated following the procedure introduced in section ‘Data pre-processing’. One-to-one sensor response reconstruction is conducted. Using these datasets, seven networks are trained and tested. The relative reconstruction error is quantified by computing and normalizing the L2-norm of the discrepancy between the reconstructed acceleration response AR and original true acceleration response AT as
The relative error is the summation of differences at all the points considered. When evaluating the error of response reconstruction in the frequency domain, AR and AT are the acceleration responses in the frequency domain obtained by fast Fourier transfer (FFT).
The coding of the DenseNets is compiled with Keras under TensorFlow backend. The computer used for training is built with a GTX2080TI GPU, an i7-6700K CPU and 16 GB memory. Each of the above-mentioned networks is trained by the training dataset for 150 epochs. The training of each network spends around 10 min and the testing can be implemented in almost real time. The training and validation losses of the DenseNets using input data from channel 15 after each epoch of training are illustrated in Figure 5. It can be observed that both the training and validation losses decrease sharply in the first several epochs because of the involved dense connection and skip connection techniques which improve the convergence speed effectively. The training loss continues to decrease with increased training epoch, while the validation loss stops decreasing around the 110 epochs and finally exceeds the training loss when 125 training epochs are finished, indicating that the network starts overfitting the training data. After the training process, a minor discrepancy between the training and validation errors is found, which means that the developed DenseNet learns robust features from the training data and also fits the validation data well. Selecting 150 epochs for training is practical, which ensures the proposed networks to be fully trained without overfitting.

The convergence curve of DenseNets with input measurement from channel 15.
The reconstruction errors for each network with different training datasets in both time and frequency domains are shown in Figure 6. The errors are computed using all the testing data in 150 min to represent more reliable results. The reconstruction errors in the time and frequency domains show similar trends. Selecting channel 1 as the input brings the largest response reconstruction error, while using channel 15 as input leads to a much smaller reconstruction error. Figures 7(a) and 8(a) demonstrate the comparison between the original true and reconstructed responses of the unavailable location at channel 17 in the time domain within an arbitrarily selected 60 s, using responses from channels 1 and 15, respectively. Figures 7(b) and 8(b) show the true and reconstructed responses of channel 17 in the frequency domain. The spectrum of the response at unavailable locations, that is, the response used for constructing the unavailable response, in the frequency domain is also shown. It should be noted that the accelerometer of channel 1 is installed at a height of 30.63 m of the GNTT. Compared to the total height of the structure which is 610 m, the vibration response of the structure under ambient excitation at the location of channel 1 is very small, and most of the vibration modes are weakly excited as shown in Figure 7(b). This leads to a low signal-to-noise ratio (SNR) and results in the reconstruction difficult based on the low-quality information. This can also explain that except reconstruction using channel 7, the reconstruction errors decrease with the rising height of the available channels with the increasing quality of the input information. Evidenced by Figures 7 and 8, it is noted that for an effective response reconstruction, the responses from the channels at available locations should contain the information of the significant vibration modes of the responses at desired locations. This idea is similar to optimal sensor placement for modal identification, which ensures that the maximum modal information is obtained and the mode shapes are highly correlated. Locations that are weakly participated in a certain vibration mode will not contribute too much valuable information for response reconstruction at the corresponding modal frequency. As a result, the measurement data from these locations may not provide essential information for reconstructing the responses related to this particular mode at other locations. In this study, for the case with 15 modes, it is impractical that all the involved modes of the desired locations are well excited at any one single node, which is far away from the desired locations. Therefore, multiple sensors are necessary for accurate dynamic response reconstruction, especially when considering a large number of vibration modes. By conducting the modal analysis of the measurement data, acceleration response from channel 7 is found to have the minimum number of overlapped modes with channel 17, which causes a relatively large reconstruction error, as observed in Figure 6. Consequently, the reconstruction performance is strongly correlated with the response information and the overlapped modes between the available channels and the desired channels.

One-to-one response reconstruction error using input measurement from different channels: (a) time domain and (b) frequency domain.

Response reconstruction of channel 17 within an arbitrarily selected 60 s using channel 1: (a) time domain and (b) frequency domain.

Response reconstruction of channel 17 within an arbitrarily selected 60 s using channel 15: (a) time domain and (b) frequency domain.
On the contrary, using input from multiple channels simplifies this issue. In this section, structural response reconstruction with multiple channels input is also conducted to reconstruct the responses at the desired locations. To test the reconstruction performance of using every single and multi-channels to reconstruct the response of channel 17, the most correlated two, three, and until all the available channels are selected and processed as the input to train different networks. The number of input channels and their corresponding channels is summarized in Table 3. The training time of multiple channels is also around 10 min for each network, since the architecture of the network and the number of training and validation samples are kept the same. Because the used dense connection and skip connection make the feature extraction and propagation very efficient, the small size of training datasets is required for constructing the nonlinear relationship. Owning to the high convergence efficiency, the training of the networks is only about 10 min, which is much shorter compared with training a traditional deep network that usually takes hours or even days. The reconstruction errors of using each trained network in time and frequency domains are shown in Figure 9. As shown, the reconstruction error decreases with the increase in available channels and achieves the minimum when data from five channels are used in the reconstruction of the response of channel 17. After that, the error increases with the number of channels used for reconstruction. It should be noted that when data from channels 1 and 7 are used, the reconstruction error is higher. This is because channel 1 is farthest from channel 17 and data from channel 7 have the minimum overlapped modes with channel 17. As discussed above, the measurement from these two channels may not provide valid information for reconstruction but slightly affect the effective training and fitting of the network model. From the deep learning aspect, guaranteeing the quality of training datasets is of great importance for promoting the network to learn effective and abstract features. For example, when conducting the image classification using CNN, mixing low-resolution or damaged photos in the training datasets will bring serious negative impact on the classification accuracy. Figure 10(a) and (b) shows the reconstructed responses using the best five channels as input to demonstrate the effectiveness of the proposed method. The responses are accurately reconstructed in the time domain with minor discrepancies, and the vibration modes are matched well in the frequency domain. The measurement data from multiple channels can provide more comprehensive information of responses and vibration modes; however, the use of low-quality data may even affect the accuracy. Nonetheless, the response reconstruction accuracy of using measurement data from multiple channels as input is much better than using a single channel, even when some interferential information from the farthest channels is involved in the input data. When selecting the optimal locations and numbers of sensors from available measurement channels for structural response reconstruction, one-to-one channel response reconstruction is conducted and sorted from the most to least correlated. The number of available channels is increased based on the order of the above-obtained sequence from results on one-to-one channel reconstruction to serve as input for multiple channel response reconstruction to choose the combination with the best performance. It should be noted that this is not the most effective way to select correct channels but is effective for most cases with a limited number of installed sensors.
The number of input channels and the corresponding channel numbers.

The reconstruction errors of using an increasing number of input channels: (a) time domain and (b) frequency domain.

Reconstruction of response at channel 17 using five channels. Comparison of (a) time domain and (b) frequency domain.
Most correlated channel is unavailable
As discussed above, the correlation between responses of available and unavailable locations is very important for structural response reconstruction. One rigorous case is when the measurement from the most correlated channel to the desired channels is not available. From Figure 6, the one-to-one reconstruction results indicate that channel 15 is the most correlated channel with channel 17 and provides the best one-to-one reconstruction accuracy. From the results as shown in Figure 9, using the measurement data from channels 1 and 7 will induce negative effects on the response reconstruction of channel 17. Therefore, the measurement data of the rest of the channels, namely, 3, 5, 11 and 13, are used as the input. The same training and testing procedures are followed as described in section ‘Data pre-processing’. To clearly show the discrepancy, a 60-s segment of the true and reconstructed responses in the time domain is plotted in Figure 11(a). The true and reconstructed responses are then transferred to the frequency domain and illustrated in Figure 11(b). The reconstructed response in the time domain shows a good agreement with the true response in both amplitudes and waveforms. Confirmed by the response in the frequency domain, the trained DenseNets accurately reconstructs all the involved vibration modes of channel 17 in both resonant frequencies and amplitudes. This indicates that the developed DenseNets successfully construct the complex nonlinear relationships between the input and output through the training process. Robust high-level features are extracted from the testing data, which is then used to accurately reconstruct the response at the desired location. By observing Figure 11(b), it can be found that the reconstructed response has a higher SNR compared to the raw measurement. This proves the advantage of using the bottleneck structure where noise components are eliminated when extracting high-level features by shrinking the feature maps. The reconstruction error is 30.11% and 23.51% for the reconstructed response in the time domain and frequency domain, respectively. The error is mainly contributed by the noise components submerged in the true response in the region of the frequency band without natural frequencies with significant energy. In contrast with the reconstruction using the measurement from the best four channels as listed in Table 3, the error is slightly higher, but the desired response can still be reconstructed with a high degree of accuracy and reliability.

Reconstruction response of channel 17 when the most correlated channel is unavailable. Comparison of (a) time domain and (b) frequency domain.
The case study to reconstruct the responses of two channels by one network is also conducted. A minor adjustment of the architecture of DenseNets is implemented, which only increases the feature maps of the final layer from two to four. As a result, the output of DenseNets becomes the measurement of two channels. The training data are generated where the input data contain measured responses from channels 3, 5, 11 and 13 and the output data consist of reconstructed responses from channels 15 and 17. Reconstructed responses for these two channels are evaluated together, where the errors for response reconstruction in time and frequency domains are 31.98% and 25.37%, respectively. With the increasing complexity of the relationships between input and output, the reconstruction error is slightly higher than reconstructing a single channel. Compared with preparing two datasets and training two networks individually, the developed network can also conduct the response reconstruction of multiple channels simultaneously, which promotes the reconstruction efficiency but slightly sacrifices the accuracy.
The noise effect
Noise is a critical issue that influences the effectiveness and accuracy of condition assessment and damage detection. The noise immunity of the proposed method is investigated by adding a certain level of white Gaussian noise to the input measurement and evaluating the accuracy of the reconstructed response. The white noise is simulated as a random vector N with the same length and sampling rate as the input response. The random vector follows a distribution with zero mean and unit standard deviation. The added noise is proportional to the root mean square (RMS) of the response A, and the noisy acceleration response
where l represents the noise level. In this study, the responses from channels 3, 5, 11, 13 and 15 are selected as the input, and the response of channel 17 is selected as the output to be reconstructed. The network is trained with the raw measurement and tested by the noisy response with 20% RMS noise added to the input of the testing data. An example of segmental original and noisy acceleration responses is shown in Figure 12. The noisy response is then inputted to the trained DenseNets, and the reconstructed response of channel 17 in time and frequency domains is demonstrated in Figure 13(a) and (b), respectively. The reconstructed response in the time domain shows a good agreement with the true one. From the frequency domain, the modal frequencies are accurately reconstructed except some minor differences in the amplitudes of a few modes. The errors of the reconstructed responses in time and frequency domains are quantified as 31.34% and 25.62%, respectively, compared with the true responses. Referring to the results in section ‘The effect of input channels’, the errors are 2.35% and 3.38% higher than using measured response without noise. The increase in reconstruction errors is moderate in term of the severity of the injected noise, demonstrating the superiority of using the proposed approach in immunizing the noise effect. The proposed DenseNets eliminate most of the noise in the higher level features, which can be attributed to the specially designed network architecture.

Comparison between the true response and noisy response with 20% noise.

Comparison of true and reconstructed responses in the (a) time and (b) frequency domain with noisy input.
Modal identification using the reconstructed response
Modal parameters including natural frequency, damping ratio and mode shape, which reflect the dynamic vibration characteristics of the structure, are widely used to develop the damage index for condition assessment and damage detection. Accurate identification of modal parameters is of significant importance for effective condition monitoring, since the identification error may mask the variation of modal parameters induced by structure condition changes. Thus, modal analysis using the true and reconstructed responses is further conducted to evaluate the applicability of the proposed method. The responses of channel 17 are reconstructed using responses of the available channels 3, 5, 11, 13 and 15. A classic operational modal identification method named frequency domain decomposition (FDD) 49 is selected for modal parameter identification using original true and reconstructed responses from channel 17. The FDD method decomposes the spectral density function matrix of the responses by singular value decomposition (SVD) and estimates the vibration modes by peak picking. The following sections will evaluate the feasibility of using reconstructed responses for modal identification, by comparing with the FDD results of true responses from qualitative and quantitative aspects.
Qualitative analysis of the reconstructed response
For providing the reference modal information, the modal parameters of GNTT are first identified using stochastic subspace identification (SSI) and FDD methods with measurements taken from all the channels as shown in Figure 14. Fifteen modes within 2 Hz are accurately identified, which agree very well with the results reported in the previous studies.47,48 By conducting the modal parameter identification with true and reconstructed responses from channel 17, FDD outputs of these two responses are shown in Figure 15(a) and (b), respectively. It can be observed from Figure 15 that modes 1, 3, 5, 8, 11, 13 and 15 are significantly excited and are accurately identified with the true and reconstruction responses of channel 17. Those modes are selected by peak picking and marked by pink points with the corresponding system mode orders. Picking the peak values from the plot of singular values is based on the engineering judgement. Consequently, the effectiveness of modal identification using FDD is dependent on the quality of responses where signals with low SNR can result in a confusing output. Observing the FDD results of using the reconstructed responses as shown in Figure 15(b), all the excited modes are effectively and clearly reconstructed without adding any extra spurious modes. It should be highlighted that the power of the response components outside the bandwidths of natural frequencies is much lower than that of natural frequencies, which reveals a high signal quality. Comparing Figure 15(b) with Figure 15(a), the SNR of the reconstructed response is better than the true response. The results demonstrate that the proposed DenseNets can effectively reconstruct the responses which are not measured or are lost and eliminate the noise components simultaneously owning to its strong capacity and specifically designed architecture. In addition, it is worth mentioning that vibration modes of GNTT are very closely spaced, which leads to the response reconstruction task more difficult for the traditional methods especially for defining the thresholds or parameters to split closely spaced modes. In contrast, the proposed method can automatically extract the hidden features from the input data and learn the complicated relationships between the input and output accurately. By analysing Figure 15(b), where the vibration modes are well reconstructed and the noise is eliminated, those learned features are very likely to be the vibrational characteristics. The results are consistent with the finding reported in an existing study 25 where the extracted high-level features from vibration responses represent the natural frequencies and mode shapes of the structure.

SSI and FDD outputs of using measurements from all channels.

Modal identification results of using (a) true response and (b) reconstructed response.
Quantitative analysis of the modal parameters
The effectiveness of the proposed method is further verified by comparing the identified modal parameters from the reconstructed responses and those from the true responses. The natural frequencies and damping ratios are identified using true and reconstructed responses, respectively. As mode shapes cannot be identified using only one channel, evaluation of the mode shapes is realized by considering the correlation between channels of true and reconstructed responses. The correlation accuracy evaluation of mode shapes is performed using the co-ordinate modal assurance criterion (COMAC), which is superior in assessing the measurement from a single DOF. 50 The COMAC of a single DOF is computed as
where m is the total number of modes and
Comparison of the identified modal parameters from true and reconstructed responses.
COMAC: co-ordinate modal assurance criterion; f: frequency; ξ: damping ratio; ϕ: mode shape.
Conclusion
This article proposes a novel approach based on DenseNets for reconstructing the acceleration responses at locations without recorded data, using recorded responses at other locations. The detailed design of the DenseNets for facilitating the reconstruction of vibration response is elaborated. The developed DenseNets can alleviate the vanishing-gradient issue, strengthen feature extraction and propagation, and substantially reduce the number of parameters, which significantly increases the training efficiency and accuracy. The effectiveness and robustness of the proposed method are demonstrated by experimental studies with real measurement data taken from GNTT. The response reconstruction is accurately performed in both time and frequency domains. Meanwhile, the identified modal parameters from the true and reconstructed responses show a good agreement. The results demonstrate that the proposed method can extract the high-level features from the input data and develop the nonlinear relationship between these features and the output responses to be reconstructed. The proposed method is also proved to have a strong noise immunity, which is practical for civil engineering applications. Further study can be conducted to investigate this topic to select the optimal sensor numbers and locations for response reconstruction using deep learning techniques.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship and/or publication of this article: The work described in this article was supported by Australian Research Council Future Fellowships (FT190100801).
