Abstract
Short-term building load forecasting is indispensable in daily operation of future intelligent/green buildings, particularly in formulating system control strategies and assessing the associated environmental impacts. Most previous research works have been focused on studying the advancement in forecasting techniques, but not as much on evaluating the availability of influential factors like the predicted weather profile in the coming hours. This article proposes an improved procedure to predict the building load 24 hours ahead, together with a backup weather profile generating method. The quality of the proposed weather profile generation model and the forecasting procedures were examined through a case study of application to university academic buildings. The results showed that the load forecasting accuracy with the application of either the real weather data on record or of the predicted weather data from the profile generation model is very much similar. This indicates that the weather prediction model is suitable for applying to building load forecasting. Besides, the comparisons between different sets of input data illustrated that the forecasting accuracy can be improved through the input data filtering and regrouping procedures.
Keywords
Introduction
Building load forecasting methods are generally categorized by the prediction time span as follows1,2: (a) long-term forecast of more than one year ahead, (b) medium-term forecast from one week ahead to one-year ahead, and (c) short-term forecast from fraction of an hour up to one week ahead. Short-term load forecasting (STLF) at the time span of 24 hours (one day) ahead is the focus of this study. This prediction time span of hourly profile can be crucial for formulating the daily operation plan of the utility systems or for smart micro-grid applications. 3 A suitable STLF will enable the utility provider to observe and take control of the balance between the supply and demand sides. Accurate STLF of a micro-grid can enhance full system integration with renewable energy resources and to improve building system efficiency and save operating costs in response to the variation of electricity price.4–7 Hence, STLF for a day ahead can be a very useful means in energy management.
In real buildings, the electricity load curves carry much variability, noise and non-linear characteristics. The occupant activities and operating schedules vary considerably with the building type, location, and perhaps more important for air-conditioned buildings, the weather conditions. 8 Numerous external and internal factors co-exist 9 and this led to the formulation of different STLF models.10–14 In order to show the sophistication in modelling techniques, many of them choose to verify the prediction accuracy by comparing the calculation results based on historical weather data (like from local observatory records) with the past building load profile (like from energy management system record). But in reality, the actual weather data of tomorrow will never be available today for future building load projection. This dilemma calls for a return to the originality15–17 of using either weather forecast data (that may be available from a third party institution) or a weather forecasting model for self-generation of the unknown data set. 18 It comes to us that the former may not be always available or the acquisition may have substantial cost implication. For this reason, we have successfully developed a STLF model that goes together with an hourly weather forecasting technique. The following presents our works in detail, which include the technological review, the influential factors analysis, the proposed forecasting framework and finally a real case demonstration.
Technological background
Influential factors and data input
The energy use profile of a building can be taken as a multi-input mathematical function. The variation depends on four main interacting factors namely building structure, operation schedule, occupancy and weather conditions. In principle, the building structure is fixed once the building project is completed. The operating schedule and occupancy (including the usage characteristics, occupancy density and activities) can be taken as the internal factors,19–28 of which the time elements are mostly included to mimic dynamic human behaviours. Many previous studies demonstrated that the historical performance at specific hours of the day and the day type could have great implication on the current hour of building energy consumption. In a university building, for example, academic calendar affects much the energy consumption pattern.
As far as STLF is concerned, the importance of individual influential factor is not the same in all building cases. The input data should be best filtered before loading for forecasting model development, and the filtering method can be either linear or nonlinear.29–32 The linear analysis methods like regression and sensitive studies may not fit the nonlinearity between the influential factors and building energy consumption rate well. In order to deal with this, Kusiak et al. 33 proposed the boosting tree algorithm as a means of input data analysis, whereas Keynia 34 recommended the data mining method based on mutual information (MI) evaluation. These nonlinear feature selection methods are effective to filter out the irrelevant (or redundant) candidate inputs.
Forecasting models
Energy prediction techniques can be grouped into three categories: the classical engineering approach (white box), the data-driven statistical approach (black box) such as artificial neuron network (ANN) and the hybrid approach (grey box). An up-to-date classification is illustrated in Figure 1.
35
With the rapid development of modern meters, the data-driven forecasting method is now widely applied, like the ANN model as an example. ANN has been successfully applied for the prediction of the overall building energy consumption and also the cooling/heating demand without prior knowledge of the building geometry or the material thermal properties.36–43
The categories of building forecasting models.
New weather forecasting technique
A new weather forecasting model to generate the upcoming 24-hour weather profile is introduced below. This simple model uses the basic weather forecast information of the next day as input, such as the daily ranges of dry-bulb temperature and relative humidity. These figures are readily accessible via public media, like television or radio broadcast or public websites. As general reference, Figure 2 shows the typical daily variation of air temperature, relative humidity and global solar radiation of Hong Kong as a subtropical coastal city. Such hourly changes of weather conditions of some well-known cities can be available on international weather forecast channels. But in other places where the future hourly profiles remain absent, our proposed forecasting model can be applied.
Typical daily weather profiles: (a) air temperature and relative humidity and (b) global solar radiation.
Temperature profile
A typical 24-hour temperature profile has a ‘peak’ (T
max
) and a ‘trough’ (T
min
) across the day. The dry-bulb temperature T of each hour can be estimated by means of the normalized temperature range (known as the multiplier value β) as follows
Derived yearly averaged 24-hour temperature profiles of Hong Kong for 2012–2015: (a) from actual observatory record and (b) normalized multiplier curves. Reference multiplier values β for application in Hong Kong.
Relative humidity
Based on the predicted ambient temperature of each hour and its definite relationship with the water vapour saturation pressure, the corresponding relative humidity can be determined from their high value (ϕ max ) and low value (ϕ min ) forecasting through the following steps, taking the fact that the water vapour pressure changes very little within a day.
Step 1: Calculate the water vapour saturation pressure from each dry-bulb temperature by
Step 2: Generate the reference water vapour pressure P
c
by
Step 3: Calculate the relative humidity as
Global solar radiation
The generation of the 24-hour global solar radiation is based on the simple sky model44,45 commonly used in solar energy simulation. On any calendar day (of day number d from 1 to 365), the time of sunrise and sunset at any position on Earth are fixed. By defining the daily angle θ as
Solar declination δ is the angle in radians between the equatorial plane and the line connecting the centres of the Sun and the Earth. This can be calculated as
The solar angular hour is calculated as
The number of hours of daylight (h
d
) is determined as a function of h, in that
Then the time of sunrise (t
sr
) and sunset (t
ss
) can be calculated as
In the simple sky model, the solar irradiance (I) is characterized by the time of sunrise and sunset, as well as the peak solar irradiance (I
max
). At any time t of the day, this is given by
By definition, I
ETI
is a function of the solar constant I0 (at 1362 W/m2), the eccentricity correction factor E
0
and the zenith angle ψ, i.e.
Derived reference coefficients of k t regression equation of Hong Kong.
Results verification
Mean absolutely precentage error of the predicted weather data against the official record of year 2015.

Weather prediction model verification: (a) temperature, (b) relative humidity and (c) global solar radiation.
New STLF model
The general framework of our STLF model is given in Figure 5, which illustrates the flow processes of data preparation, filtering, regrouping and then the forecast model structure optimization. At first, based on the full list of anticipated influential factors, the raw input data set is sorted for analysis. Then, the importance of these influential factors is ranked. The less important factors will be filtered out, subject to the screening criteria. Afterwards, the input data on calendar-day 24-hour basis are regrouped into sub-sets based on the distinct energy consumption characteristics. For illustration convenience, the decision of forming two data sub-sets as the result is assumed in Figure 5. Then for each data set, the forecasting model structure is optimized through error analysis. The above-described stages are further elaborated below:
(a) Data filtering: Filtering the input data is needed in order to overcome the curse of dimensions which caused by irrelevant and redundant candidate inputs. So before the confirmation of input variables, the input data are to be analysed using the MI criterion. The involved influential factors will have the individual MI values calculated then their importance ranked. The detailed algorithm can be found in Keynia.
34
The overview of this data filtering process can be seen in Figure 6. (b) Data regroup: Building energy consumption usually has its case-specific trend characteristic. It will have similar daily consumption profile during the same operation schedule, independent of the weather condition. Based on the building operation schedule and the energy consumption outcome, the input data can be regrouped into sub-sets consisted of 24-hour daily data as the group members. Then, a distinct forecasting model will be worked out for each sub-set. This modelling procedure will lead to more accurate predictions. The details of this part of data handling can be referred to our case study in section ‘Case study’. (c) Forecasting model and structure optimization: Our forecasting model is based on ANN model with back-propagation and Levenberg-Marquardt learning method. Figure 7 shows an example of the ANN architecture. In order to achieve high computation speed, only one hidden layer was adopted. As the results accuracy can be sensitive to the transfer function and the neuron numbers of the hidden layer,
31
systemic comparison runs should be performed in the model structure optimization process, which can be seen in the section ‘Energy consumption analysis and data regrouping’. (d) Accuracy evaluation criteria: The following three accuracy criteria are commonly used to assess forecasting model effectiveness: the MAPE, the Daily Peak MAPE, and the coefficient of variance of the root mean squared error (CVRMSE) under 95% confidence limits.
47
By definition
Forecasting flowing chart. Overview of data filtering process. ANN forecasting model architecture. ANN: artificial neuron network.



Case study
A case study is introduced in this section as an illustrating example. This case study involved two multi-story academic buildings in our university campus. Building No. 1 was an older building built over 20 years ago. Building No. 2 was relatively new, of age less than five years. Energy is supplied from the mains electrical grid. The two buildings were fully air conditioned with year-round space cooling demand. Their daily operation schedules were the same, with opening hours from 7:00 to 23:00 on normal weekdays and 7:00 to 18:00 during weekends. The university operation was according to the academic calendar with four distinct period types: teaching semester, student revision week, examination period and semester break.
Data preparation and filtering
From the past years record, the available forecast/actual weather data, historical energy consumption record and academic calendar period type of each day were gathered for formulating the forecast model. They were used as the input data to predict the 24-hour energy consumption profile of the following day. Weather data of 2013 and 2014 were first collected from the Hong Kong Observatory. These include dry-bulb temperature, relative humidity, global solar radiation, rainfall, clearness of sky, cloud condition and wind speed. The energy consumption data were the hourly electricity consumption record of each building extracted from the central energy management system. Influential factors were identified based on the findings of the previous works and also the available data at hand.
Energy consumption influential factor analysis.
MI: mutual information.
Comparison of mean absolutely percentage error in energy prediction – with or without wind speed.
Energy consumption analysis and data regrouping
Before performing the forecast model evaluation, the energy consumption behaviour of these two buildings in 2013 and 2014 was analysed. Figures 8(a) to (f), respectively, show the load cloud charts of Buildings No. 1 and No. 2 and for three calendar months with seasonal representation, i.e. August 2013 for summer, April 2014 for spring and December 2014 for winter. In these charts, the x-axis shows the calendar dates and the y-axis shows the daily hours 1–24.
Energy consumption (kW) cloud charts of Buildings No. 1 and No. 2. (a) Building No. 1 in August 2013, (b) Building No. 2 in August 2013, (c) Building No. 1 in April 2014, (d) Building No. 2 in April 2014, (e) Building No. 1 in December 2014 and (f) Building No. 2 in December 2014.
The daily profile and weekly trend can be readily observed from the colour distribution patterns, in that, the situations of both buildings look alike. In general, at night from 00:00 AM to 8:00 AM, the energy consumption level was very low. Then after 8:00 AM, the energy consumption increases rapidly, reaching the peak level in between 14:00 PM and 17:00 PM. Then after 20:00 PM, the energy consumption decreases.
Testing mean absolutely percentage error of different transfer function combinations.
Testing mean absolutely percentage error of different neurons number.
Forecasting results and error analysis
The rationale of our proposed methodology was tested by comparing the predicted energy consumption with the actual records of 2015, and also the error analysis. The results of comparison and analysis can be found in Figure 9 and Figure 10, Tables 8 and 9. Our comparisons show that the results of forecasting agree well with the real load consumption, as reflected in Figures 9(a) and (b) for Buildings No. 1 and No. 2, respectively. For each building, two sets of prediction results were acquired with the use of the same energy forecasting model. The difference was in their weather data input, in that, data from the weather forecast model were applied in the first trial, but actual hourly weather data of the day were used in the second trial and then compared. It was found that the prediction accuracy was very good for both trials. Figure 9 shows the graphical plots for a summer week. Both load prediction curves show comparable performance with the actual consumption. The forecasting error distribution has good agreement with normal distribution, seen in Figure 10, which shows the reasonable of prediction results.
Results comparison of load forecasting for the week 8–14 July 2015 using different weather data sets: (a) Building No. 1 and (b) Building No. 2. Forecasting error distribution analysis. (a) Building 1: Prediction with record weather data, (b) Building 1: Prediction with forecast weather data, (3) Building 2: Prediction with record weather data and (d) Building 2: Prediction with forecast weather data. Comparison of mean absolutely percentage error in energy prediction – actual weather data against forecast. Forecasting errors with or without input data filtering and regrouping. CVRMSE: coefficient of variance of the root mean squared error; MAPE: mean absolute percentage error.

Also from Table 8, it can be concluded that the MAPE differences (of only less than 0.2%) of the predicted energy consumption based on the weather data obtained from the forecast model and from the real weather data are very small. This so happens to both buildings. The quality of the weather profile prediction technique is therefore fully demonstrated. Generally speaking, the modelling accuracy of Building No. 1 is better than Building No. 2. This is because the occupancy in the older building is more stable than in the new Building No. 2.
The effectiveness of having filtering and regrouping in this case study was evaluated making use of the Building No.1 performance. The results of forecasting errors are shown in Table 9. It can be seen that for both Set 1 and Set 2, the prediction accuracy can be improved with filtering. The use of more influential factors (without filter) was not helpful at all. So, the filtering process is proved appropriate. The same conclusion can be reached with the use of regrouping, as demonstrated by the error analysis with the use of the same evaluation parameters. It is worth mentioning that the MAPE of 5.98% in this case study with filtering is generally lower than those quoted in the previous related studies, like those in Bagnasco et al., 31 Dong et al., 41 Yuce et al. 42 and Marmaras et al. 43 are 7.0%, 9.26%, 6.51% and 6.7%, respectively.
Conclusion
This paper proposed a framework to predict the short-term building load with data filtering and regrouping which will improve the forecasting accuracy. It also provides a backup weather forecasting method to generate the needed weather data used in the load forecasting model. The main conclusion of this paper is listed below:
(1) The input data filtering procedures show that dry-bulb temperature, relative humidity and global solar radiation were important influencing factors, so was the energy consumption level of the same hour one day before. (2) Observations through the cloud charts identified the necessity of data regrouping based on normal weekday and the remaining periods of the academic calendar. The subsequent error analysis confirms the effectiveness of input data filtering and regrouping to improve the forecasting model accuracy. (3) The proposed weather forecasting method could be used as backup approach to generate the needed data for the load prediction model, especially for the case without large variation of occupancy and climate. In order to extend the adaptivity of the proposed model, more case study need to be consider in our future work.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The work described in this article was financially supported by the Contract Research Grants of the City University of Hong Kong (Project no. 9231136).
