Int. Renew. Energy Dev. 2025, 14 . , 1221-1234 | 1221 Contents list available at CBIORE journal website International Journal of Renewable Energy Development Journal homepage: https://ijred. Research Article Multi-objective HVAC control using genetic programming for gridresponsive commercial buildings Sibtain Waheed* and Shuhong Li School of Energy & Environment. Southeast University. Nanjing. China 211189. China Abstract. Commercial buildings are significant energy consumers, with their heating, ventilation, and air conditioning (HVAC) systems being major Optimizing these systems is crucial for energy conservation, yet advanced artificial intelligence methods like Deep Reinforcement Learning (DRL) often produce opaque black-box solutions. While post-hoc explanation methods can offer some insight, they are often inexact and fail to render the core decision logic fully transparent, hindering trust and practical implementation. This paper presents a novel approach using Genetic Programming (GP) to automatically design HVAC control strategies that are both highly effective and inherently understandable. The novelty of our framework lies in its direct evolution of interpretable, multi-objective control policies that holistically co-optimize energy efficiency, occupant thermal comfort, and integrated Demand Response (DR) for a complex multi-zone system a combination not extensively explored in prior GP-HVAC We applied this framework to manage the central air handling unit of a simulated multi-zone office building, enabling it to dynamically adjust key settings like air temperature and fan pressure. Rigorous testing in a validated EnergyPlus simulation environment showed that the GPdesigned control policies reduced annual HVAC energy use by 40. 9% compared to standard ASHRAE A2006 guidelines, 28. 4% against the advanced ASHRAE G36 standard, and a notable 9. 3% more than a state-of-the-art DRL controller. These substantial energy savings were achieved while maintaining excellent occupant thermal comfort for 98. 8% of occupied hours. Furthermore, the GP controller demonstrated robust performance during Demand Response scenarios, achieving a 72. 1% reduction in peak power draw. A key outcome is that these high-performing strategies are expressed in a transparent format allowing direct inspection and understanding. This research establishes Genetic Programming as a compelling method for creating intelligent HVAC controls that are not only efficient and grid-responsive but also transparent, fostering greater confidence in advanced building automation. Keywords: Genetic Programming. HVAC Control. Energy Efficiency. Transparent Control. Demand Reponses. Building Automation @ The author. Published by CBIORE. This is an open access article under the CC BY-SA license . ttp://creativecommons. org/licenses/by-sa/4. 0/). Received: 20th May 2025. Revised: 7th Sept 2025. Accepted: 4th Oct 2025. Available online: 19th Oct 2025 INTRODUCTION Commercial buildings use a lot of energy, and their heating, ventilation, and air conditioning (HVAC) systems are often the main reason, accounting for about 40-50% of total building energy consumption in many cases (Ghaderian & Veysi, 2021. Kaushik et al. , 2. With growing concerns about climate change, rising energy costs, and the need for smarter power grids, making these systems more efficient is a big deal. Traditional HVAC controls, like those based on fixed rules from standards such as ASHRAE 90. 1, work fine but often miss opportunities to save energy because they don't adapt well to changing conditions like weather, occupancy, or peak demand times (Pyrez-Lombard et al. , 2008. Amer et al. , 2024. Yoon et al. This can lead to wasted energy, uncomfortable indoor spaces, and higher bills. Over the years, researchers have turned to advanced methods to make HVAC smarter. Model predictive control (MPC) use math models to predict and adjust settings ahead of time, which can cut energy use by 20-30% in some studies (Afroz et al. , 2018. Bitar et al. , 2. Then there's artificial intelligence, especially reinforcement learning (RL), where systems learn from trial and error to balance energy savings with comfort (Xie. Ajagekar, & You, 2023. Al Sayed et al. , 2. Deep reinforcement learning (DRL) takes this further by handling complex data, showing promising results in simulations and even real buildings (Lu et al. , 2022. Sanzana et , 2. But the problem with many AI methods is they're "black-box" models. You get great performance, but it's hard to understand why the system makes certain decisions. This lack of transparency can make building managers hesitant to trust and implement them, especially in critical setups where safety and reliability matter (Pinto et al. , 2022. Pinthurat. Surinkaew, & Hredzak, 2. To address these limitations, advanced artificial intelligence (AI) techniques have gained prominence. Deep Reinforcement Learning (DRL) has emerged as a leading model-free approach, demonstrating 15-30% energy savings over conventional controls in simulated multi-zone buildings (Yu et al. , 2021. Hou et al. , 2. Algorithms like Deep Q-Networks (DQN). Soft Actor-Critic (SAC), and multi-agent variants enable adaptive policies that learn from interactions with the environment, optimizing actions such as SAT and DSP setpoints (Kumar et al. Niazi et al. , 2025. Sun et al. , 2. Recent studies have * Corresponding author Email: sibtainwaheed@seu. cn (S. Wahee. https://doi. org/10. 61435/ijred. ISSN: 2252-4940/A 2025. The Author. Published by CBIORE S. Waheed and S. Int. Renew. Energy Dev 2025, 14. , 1221-1234 | 1222 extended DRL to incorporate DR, achieving 40-45% peak load reductions by preemptively adjusting HVAC operations during high-price signals (Kumar et al. , 2025. Niazi et al. , 2025. yNinar & Abut, 2. However. DRL's reliance on opaque neural networks poses a major barrier: the "black-box" nature hinders interpretability, trust, and deployment in safety-critical systems (Kargar & Bahamin, 2. Efforts to enhance transparency through post-hoc Explainable AI (XAI) methods, such as SHAP (SHapley Additive exPlanation. and LIME (Local Interpretable Model-agnostic Explanation. , have provided partial insights into feature importance (Mariano-Hernyndez et al. , 2021. Yao et , 2. Yet, these approximations often fail to capture holistic decision logic, leading to incomplete or unreliable explanations (Yan et al. , 2. Parallel advancements in model-based optimization, particularly Model Predictive Control (MPC), offer structured MPC uses physics-based models to forecast and optimize HVAC operations, reporting up to 25% energy savings while integrating DR (Bouabdallaoui et al. , 2021. Tomys. Lymmle, & Pfafferott, 2. Hybrid approaches combining MPC with machine learning further improve robustness to uncertainties (Alimohammadisagvand. Jokisalo, & Siryn, 2. However, developing accurate models is resource-intensive, and scalability remains a challenge for large buildings (Chaturvedi. Rajasekar, & Natarajan, 2020. Bouabdallaoui et al. In the DR domain, rule-based strategies provide transparency but often compromise comfort during load shedding (Cheraghi & Jahangir, 2023. Cho. Lee, & Heo, 2023. Choi et al. , 2. DRL-enhanced DR, while effective, inherits interpretability issues (Ding. Cerpa, & Du, 2. Emerging VPP concepts aggregate buildings with renewables but require interpretable controls for market participation (Pang et al. Genetic Programming (GP), an evolutionary computation technique, addresses these gaps by evolving interpretable expression trees or decision rules directly from data (Cpalka. Aapa, & PrzybyC, 2018. Sipper & Moore, 2. Prior GP applications in buildings include single-objective optimizations, such as chiller sequencing . ielding 10-20% saving. or thermal comfort in passive designs (Gao et al. , 2020. Es-sakali et al. Multi-objective GP using NSGA-II has optimized energy and comfort in residential settings (Pang et al. , 2. However. GP has not been extensively applied to integrated multi-zone HVAC control with DR, nor benchmarked against DRL in renewable-integrated grids. This study bridges these gaps by proposing a GP framework to evolve transparent, multi-objective policies for AHU control in grid-responsive buildings. Key contributions Development of a comprehensive GP framework for evolving interpretable multi-objective HVAC control policies with integrated Demand Response Rigorous evaluation of the evolved GP controllers against both conventional rule-based approaches (ASHRAE 2006 and Guideline . and state-of-the-art DRL controllers in a validated EnergyPlus simulation Analysis of the evolved control strategies, revealing how GP discovers sophisticated yet transparent operational patterns that effectively balance energy efficiency, comfort, and grid responsiveness. Validation of GP as a compelling approach for creating intelligent building controls that are not only efficient and grid-responsive but also transparent and The remainder of this paper is organized as follows: Section 2 details the methodology, including the benchmark building scenario, simulation environment. GP Figure framework, and baseline controllers. Section 3 presents the results and discussion, analyzing the evolutionary process, comparative performance, and characteristics of the evolved policies. followed by concluding remarks in Section 4. Methodology This section demonstrates the comprehensive methodology developed and employed for the direct evolution, simulation, and rigorous evaluation of interpretable, multi-objective HVAC control policies using Genetic Programming (GP), with a specific focus on integrated Demand Response (DR) capabilities. The overall research process is visually summarized in Fig. We commence by detailing the benchmark building scenario and the high-fidelity simulation environment. Subsequently, the GP-based control strategy formulation is presented, encompassing its representation, multi-objective fitness evaluation, and evolutionary algorithm configuration. describe the implementation of state-of-the-art Deep Reinforcement Learning (DRL) and established ASHRAE standard controllers, which serve as baselines for comparative analysis, along with the key performance indicators used for Benchmark Scenario and Simulation Environment To ensure a realistic and challenging testbed for the control algorithms, a high-fidelity simulation environment was meticulously constructed. This environment leverages EnergyPlus (Version 9. for dynamic building thermal and HVAC system simulation. The control algorithms, including the novel GP framework and baseline controllers, were implemented in Python (Version 3. , interfacing with EnergyPlus via the Functional Mock-up Interface (FMI) Building Model The architectural testbed is a representative five-zone Fig. 1 Methodological Framework ISSN: 2252-4940/A 2025. The Author. Published by CBIORE S. Waheed and S. Int. Renew. Energy Dev 2025, 14. , 1221-1234 | 1223 Fig. 2 Building Schematic diagram thermodynamically adapted from the U. Department of Energy (DOE) Commercial Reference Building specification for a Medium Office (New Construction. Post-1980. Climate Zone 5A equivalen. (Yu et al. , 2. The model features a total conditioned floor area of approximately 550 mA, distributed across one thermally distinct interior core zone and four perimeter zones (North. East. South, and Wes. , each subject to varying solar exposures and envelope heat transfer A schematic plan view illustrating the building layout and zonal configuration is presented in Fig. Detailed construction assemblies for walls, roof, floor, and fenestration, along with their respective thermal properties (U-values. Rvalues, thermal mass characteristic. , adhere to the reference building specifications designed to meet ASHRAE Standard Air infiltration rates are modeled based on ASHRAE standards for air changes per hour (ACH). HVAC System Model The building model is equipped with a centralized Variable Air Volume (VAV) Air Handling Unit (AHU) that conditions and distributes air to the five thermal zones. A high-level overview illustrating the primary functional blocks of the HVAC system and the main interaction points with the Genetic Programming (GP) controller is provided in Fig. Delving into the specifics of the air-side system, the AHU, whose internal schematic and detailed GP control intervention points are depicted in Fig. 4, comprises an outdoor air economizer section, a chilled water cooling coil, and a variablespeed supply fan. The economizer operation is based on a drybulb temperature comparison between outdoor and return air, with minimum outdoor air ventilation rates continuously maintained according to ASHRAE Standard 62. 1-2019 during occupied periods. The cooling coil is supplied with chilled water from an electric water-cooled chiller plant. This plant includes the primary chiller unit, an open-loop cooling tower for heat rejection, and associated variable-speed primary and secondary chilled water pumps, as well as condenser water pumps. The operational logic and setpoints for this chiller plant, particularly the chiller supply water temperature, are influenced by the evolved GP policies as detailed in Section 2. Air distribution to the conditioned spaces is managed by five pressureindependent VAV terminal units, one serving each thermal For baseline operations and scenarios within this study, the VAV terminal units modulate their dampers to meet zonespecific cooling demands based on standard ASHRAE 2006 control sequences. The GP controller focuses on optimizing the Fig. 3 High-Level HVAC System Overview with GP Controller Interaction ISSN: 2252-4940/A 2025. The Author. Published by CBIORE S. Waheed and S. Int. Renew. Energy Dev 2025, 14. , 1221-1234 | 1224 Fig. 5 Detailed AHU Schematic with GP Control Points central AHU and plant operation, while the VAV boxes react to the supplied air conditions and zone loads according to this established baseline logic. The performance characteristics fan curves, coil capacity curves, chiller COP curve. and nominal design parameters for all primary HVAC components are modeled using standard EnergyPlus objects and empiricallyderived performance curves representative of typical commercial equipment. Specific details regarding these component models are comprehensively tabulated in Appendix Table A1. Operational Conditions and Location All simulations were conducted for a continuous one-month period representative of a significant cooling season, specifically July. The meteorological conditions are based on a Typical Meteorological Year 3 (TMY. weather file for Turin. Italy (EPW file: ITA_Torino-Caselle. 160590_TMYx. , providing hourly data for temperature, humidity, solar radiation, and wind Standardized commercial office operational schedules were implemented for weekday (Monday-Frida. :00-19:00 peak, with scheduled ramp-up/dow. , lighting . ased on ASHRAE 90. 1-2019 Lighting Power Density allowance. , and miscellaneous equipment loads . lug load. Weekend operation assumes an unoccupied building state. The occupied period thermostat cooling setpoint for all zones is maintained at 24AC, with a heating setpoint of 21AC . hough active heating demand is minimal during the simulated July The central HVAC system operates from 06:00 to 19:00 on weekdays, allowing for a pre-cooling period before nominal occupancy begins. To evaluate DR capabilities, a simulated DR program was integrated. This program introduces a high electricity price signal, three times the baseload electricity price, during peak demand hours . :00-17:. on selected high-load weekdays within the simulation month. This DR signal serves as an explicit input to the intelligent controllers (GP and DRL), prompting them to modulate HVAC energy consumption. Genetic Programming (GP) Framework for HVAC Control This research proposes a GP framework to directly evolve interpretable, multi-objective control policies for the AHU. The GP aims to identify policies that holistically optimize energy efficiency, occupant thermal comfort, and effective participation in DR events. The iterative interaction of the GP algorithm with Fig. 4 GP Interaction with Simulation Environment the EnergyPlus simulation environment during the evolutionary process is conceptually illustrated in Fig. Problem Formulation for GP The GP-evolved policies are responsible for determining the following primary AHU control outputs at each discrete control timestep. OIyei . et to 15 minute. ISSN: 2252-4940/A 2025. The Author. Published by CBIORE S. Waheed and S. Int. Renew. Energy Dev 2025, 14. , 1221-1234 | 1225 Fig. 6 Evolved GP Policy Tree AHU Supply Air Temperature Setpoint ycNycIyaycN,ycycy. ,: The target temperature for air leaving the AHU cooling AHU Duct Static Pressure Setpoint ycEyaycE,ycycy . : The target static pressure to be maintained in the main supply duct by the variable-speed fan. Chiller Supply Water Temperature Setpoint ycNycIycOycN,ycycy . : The target temperature for water leaving the chiller. Economizer Control Parameter ycyceycaycuycuycu . : A parameter guiding economizer operation, such as a maximum allowable mixed air temperature or a direct outdoor air fraction command, depending on the chosen GP representation. The set of input variables available to the GP for constructing these control policies encompasses real-time and predicted environmental conditions, building thermal states, and operational signals. A comprehensive list of all input terminals and control outputs, along with their descriptions and units, is provided in Appendix Table B. and include current outdoor air temperature ycNycCyaycN , predicted outdoor air temperatures for the next . , 2, and . cNycCyaycN,ycyycyceycc 1Ea , etc. ), average and maximum zone air temperatures ( ycNycycuycuyce,ycaycyci ycNycycuycuyce,ycoycaycu ), ( yuuycNycycuycuyce,ycoycaycu_yccyceyc ), current supply air temperature ( ycNycIyaycN,ycaycayc ), return air temperature . cNycIyceyc ), an indicator of maximum VAV box damper position across zones . cOyaycOyccycaycoycy,ycaycycnyc ), current time of day, day of the week, and the active Demand Response signal . cIyaycI . GP Representation The GP representation uses a tree-based structure where each control output is determined by a separate expression tree as shown in Fig. These trees combine input variables . through mathematical operations . ddition, subtraction, multiplication, divisio. , comparison operators . reater than, less than, equal t. , and conditional logic . f-then-els. This representation allows for the evolution of complex, non-linear control strategies while maintaining human interpretability. Fitness Evaluation and Multi-Objective Optimization The fitness of each GP individual evolved policy is evaluated based on its performance over the simulation period one representative week in July including DR events, or the full The operational interaction between a candidate GP policy and the simulated EnergyPlus environment over a typical 24-hour cycle is depicted in Fig. 7, the Environment light green bar represents the continuous building simulation. During Fixed Action Periods . ours 00:00-06:00 for night setback and 19:0023:00 for system of. , pre-defined control actions are applied. During the GP Control Active period . ight blue bar, 06:00-19:00, indicated by green vertical lines on the timelin. , the evolved GP policy is engaged. At each control timestep within this active period, the GP policy receives Observations (Env. to GP) . rown downward arrow. from the EnergyPlus environment and, based on its evolved logic, issues Actions (GP to Env. ) . lue upward arrow. back to the simulation to control the AHU. The performance . nergy, comfort. DR KPI. resulting from these actions over the evaluation period ycNyceycycayco is then used to calculate the objective function values: yayaycuyceycyciyc , yayaycuycoyceycuycyc and yayaycI . Integrated HVAC Energy Consumption . ayaycuyceycyciyc ): This objective from equation . reflects the total electrical energy consumed by the chiller plant . hiller, cooling tower fans, pump. , the AHU supply fan, and any auxiliary HVAC components over the evaluation period ycNyceycycayco . J Energy = E Teval chiller_plant . , uGP . ), x . )) Pfan . , uGP . ), x . )) ) dt Here, ycEycaEaycnycoycoyceyc_ycyycoycaycuyc . and ycEyceycaycu . are the instantaneous power demands . W) of the chiller plant and AHU fan, respectively. These are functions of the GP-dictated control actions ycOyaycE . and the overall system state ycu. The integral is numerically approximated from the discrete-time simulation outputs. Aggregated Thermal Discomfort ( yayaycuycoyceycuycyc ): This objective from equation . quantifies occupant dissatisfaction due to deviations from the defined thermal comfort band. It is formulated as the sum of Time-Integrated Absolute Error (TIAE) or Integrated Squared Error (ISE) of zone temperatures from the comfort band limits . cNyaya,ycn = ycNycyceycycyycuycnycuyc,ycn Oe yuycNycaycuycoyceycuycyc and ISSN: 2252-4940/A 2025. The Author. Published by CBIORE S. Waheed and S. Int. Renew. Energy Dev 2025, 14. , 1221-1234 | 1226 Fig. 7 GP Interaction with Simulation Environment ycNycOya,ycn = ycNycyceycycyycuycnycuyc,ycn yuycNycaycuycoyceycuycyc for zone i, with ycNycyceycycyycuycnycuyc,ycn = 24Oo ya and yuycNycaycuycoyceycuycyc = 1Oo ya during occupied hours Occi . ) . J Comfort = Eu i =1E Teval Occi . ) E E max. Tzone ,i . , uGP . ), x. )) Oe TUL ,i ) E EE dt E max. TLL ,i Oe Tzone ,i . , uGP . ), x. ))) E (AC E hr ) The evolutionary search for optimal control policies is guided by a multi-objective fitness function, designed to concurrently satisfy performance criteria related to energy consumption, thermal comfort, and DR participation. Each candidate GP policy is evaluated by deploying it in the Energy Plus simulation for a representative period ycNyceycycayco . arefully selected week in July encompassing diverse load conditions and DR event. The performance is quantified by the following three objective functions, all of which are to be minimized: Demand Response Performance Deficit ( yayaycI ): This objective penalizes the failure to meet DR targets or maximizes the benefit from DR participation in equation . One formulation is to minimize the average HVAC power consumption during DR event periods, ycNyaycI , relative to a baseline power consumption, ycEycaycaycyceycoycnycuyce,yaycI . , or to penalize exceeding a DR power cap ycEycaycaycy,yaycI: J DR = PHVAC ,total . , uGP . ), x. ))dt . W) O TDR O E tEaTDR J DR = E tEaTDR max. PHVAC ,total . , uGP . ), x . )) Oe Pcap . DR )dt . To address these multiple, often conflicting, objectives, the Nondominated Sorting Genetic Algorithm II (NSGA-II) (Ghaderian & Veysi, 2. is employed as the evolutionary search NSGA-II identifies a set of Pareto-optimal solutions, representing the best achievable trade-offs between yayaycuyceycyciyc , yayaycuycoyceycuycyc , and yayaycI . From this final Pareto front, a single representative GP policy was selected for the detailed comparative analysis against the baseline controllers. This selection was based on a balanced performance criterion, specifically targeting the solution that minimized the Euclidean distance to the utopia point . , 0, . in the normalized threeobjective space. This approach was chosen to identify the policy offering the most compelling compromise across all three objectives, rather than a policy that might excel in one objective at the significant expense of others. Furthermore, to address the potential for overfitting, it is important to note that for this study, both the evolution and final evaluation of the GP policies were conducted on the representative month of July. This approach was deliberately chosen to create a controlled and reproducible 'level playing field' for directly comparing the learning capabilities of GP against the DRL agent, ensuring both were optimized under identical conditions. The critical question of generalization to unseen weather data is a primary focus for future work, as discussed in Section 4. Evolutionary Algorithm Configuration The GP system was implemented using the DEAP (Distributed Evolutionary Algorithms in Pytho. Version 1. Key evolutionary parameters were set as follows: population size of 200, number of generations set to 100, tournament selection with tournament size of 3. The entire evolution process, conducted on an Intel Core i9-12900K CPU and 32GB RAM, took approximately 30 hours to complete. Genetic operators included one-point subtree crossover with probability 0. 7 and subtree mutation with probability 0. Additionally, point mutation for ephemeral random constants was applied with a probability of 0. The initial population was generated using the ramped half-and-half method with tree depths ranging from 2 to 6, and a maximum tree depth of 12 was enforced to prevent excessive bloating. Elitism, preserving the top 2% nondominated solutions, was incorporated to ensure convergence towards high-quality solutions. These parameters were chosen based on a combination of established practices in GP literature for complex control problems (Cpalka. Aapa & PrzybyC, 2. and a series of preliminary tuning experiments. The population size and number of generations were selected to provide a sufficient search diversity and convergence time, while remaining within feasible computational limits for the highfidelity simulation environment. The crossover and mutation rates were set to standard values that encourage a balance between exploration of new solutions and exploitation of highperforming genetic material. Baseline Controller Implementations For a comprehensive performance benchmark, the evolved GP policies were compared against three distinct baseline controllers, simulated under identical environmental and operational conditions. ISSN: 2252-4940/A 2025. The Author. Published by CBIORE S. Waheed and S. Int. Renew. Energy Dev 2025, 14. , 1221-1234 | 1227 Table 1 Comprehensive Performance Comparison of Control Strategies Performance Indicator Unit Evolved GP Policy Energy Efficiency Total HVAC Energy Consumption 6,800 DRL (SAC) Baseline ASHRAE G36 Baseline ASHRAE A2006 Baseline 7,500 9,500 11,500 Chiller Plant Energy Consumption 4,080 4,500 5,800 7,000 AHU Fan Energy Consumption 1,840 2,025 2,470 3,000 % Savings vs. A2006 (Tota. Ai % Savings vs. G36 (Tota. Ai 1% (Wors. Thermal Comfort ZAT Violation Degree-Hours ACAhr Occupied Hours in Comfort Band Average Peak Load Reduction % Peak Load Reduction 5%* 1%* Total Energy Saved/Shifted Demand Response Effectiveness Deep Reinforcement Learning (DRL) Baseline same computational hardware. Further details on DRL hyperparameters . earning rates, discount factor , target smoothing E, entropy coefficient ) are provided in Appendix C. 1 (DRL Agent Hyperparameter. A DRL agent based on the Soft Actor-Critic (SAC) algorithm was developed as a state-of-the-art learning-based benchmark. State and Action Spaces: The DRL agent utilized the same state observation space as the GP terminals (Appendix Table B. Its continuous action space corresponded to AHU . cNycIyaycN,ycycy , ycEyaycE,ycycy , ycNycIycOycN,ycycy , ycyceycaycuycuycu ), normalized to [-1, . and subsequently scaled to physical operational limits. Reward Function: The instantaneous reward ycIyc as shown in equation . was formulated to align with the GP's multi-objective nature, typically as a negatively weighted sum of penalties reflecting energy consumption . cEyaycOyaya,yc ), thermal discomfort . aycaycuycoyceycuycyc,yc ), and DR non-compliance . ayaycI,yc ): E wE E PHVAC ,t E Rt = Oe E wC E Eu i =1Occi . ) E Dcomfort ,i ,t 1 E EE wDR E S DR . ) E DDR ,t EE The terms comfort I ,t 1 and ycycyc,yei are analogous to the integrands in ycycyeayeayeNyeayeeyei and ycycyc respectively, evaluated at time t or t 1. The weights . eoyc , yeoyc , yeoycyc ) were determined through an iterative tuning process. A grid search over a range of plausible weight ratios was conducted in short, preliminary training runs . k timesteps eac. The set of weights that demonstrated the most stable learning curve and resulted in agents achieving a balanced performance across energy, comfort, and DR objectives during these initial runs was selected for the full, long-duration training. Network Architecture and Training: Both actor and critic networks in SAC were implemented as fully connected multi-layer perceptions (MLP. with two hidden layers of 256 neurons each, employing ReLU activation functions, and appropriate output activations (Tanh for bounded action. The agent was trained for [ 2 C 106 simulation timestep. using an experience replay buffer of size [ 105 ]. This training process took approximately 48 hours on the ASHRAE Guideline 36 Baseline The advanced rule-based control sequences specified in ASHRAE Guideline 36-2021 (Yoon et al. , 2. were implemented as a high-performance conventional baseline. This Dynamic supply air temperature (SAT) reset using the Trim & Respond algorithm, responsive to aggregate zone cooling demand and outdoor air temperature. Dynamic duct static pressure (DP) reset using Trim & Respond logic based on VAV terminal damper positions, thereby minimizing fan energy while ensuring terminal Economizer control based on differential dry-bulb or enthalpy comparison, with enforcement of minimum ventilation requirements. Chiller plant sequencing and chilled water temperature reset consistent with standard Guideline 36 specifications. As illustrated in Fig. 8a, the SAT and DP reset logic ensures supervisory-level efficiency by adjusting setpoints dynamically in response to load conditions and zone demands. A more conventional rule-based controller consistent with ASHRAE 2006 sequences was also considered as a fundamental baseline. This configuration features a simpler SAT linear reset tied directly to outdoor air temperature, potentially fixed or less adaptive DP setpoints, and standard economizer control. The VAV terminal unit control logic . hown in Fig. was applied consistently across all baseline scenarios, including those served by the GP-controlled AHU. Each pressure-independent VAV box modulated damper position to maintain the local zone temperature setpoint, subject to minimum ventilation Performance Evaluation Metrics The comparative performance of the evolved Genetic Programming (GP) policies and the baseline controllers was ISSN: 2252-4940/A 2025. The Author. Published by CBIORE S. Waheed and S. Int. Renew. Energy Dev 2025, 14. , 1221-1234 | 1228 Fig. 8 ASHRAE Guideline 36 SAT and VAV Box Control Logic quantitatively assessed using a suite of Key Performance Indicators (KPI. , aggregated over the entire simulation period (Jul. These indicators were selected to provide a holistic evaluation across energy efficiency, occupant thermal comfort. Demand Response (DR) effectiveness, and controller Total HVAC Energy Consumption . ayaycOyaya,ycycuycycayco ): This primary energy KPI from equation . represents the sum of all electrical energy . n kW. consumed by the HVAC system components, including the chiller plant . hiller, cooling tower fans, pump. and the Air Handling Unit (AHU) fan, over the simulation period Tsim : EHVAC ,total = E Tsim chiller_plant . ) Pfan . ) Paux_pumps . ) ) dt where ycEycaEaycnycoycoyceyc_ycyycoycaycuyc . , ycEyceycaycu . and ycEycaycycu_ycyycycoycyyc . are the instantaneous power demands of the respective components. Thermal Comfort - ZAT Violation Degree-Hours ( ycOyayaycsyaycN ): This metric quantifies the integrated magnitude and duration of thermal discomfort in equation . It is calculated as the sum of absolute temperature deviations outside the defined comfort band . AC - 25AC) during occupied hours . cCycaycaycn . = . for all ycAyc zones: E max. Tzone,i . ) Oe TUL ,i ) E N z Tsim VDH ZAT = Eu i =1E Occi . ) E E E max. T Oe T . )) EE LL ,i zone ,i E Thermal Comfort - Percentage of Occupied Hours within Comfort Band (%ycNycaycuycoyceycuycyc ): This provides an intuitive measure of comfort, representing the percentage of total occupied hours during which all monitored zone temperatures were maintained within the . AC - 25AC] comfort band. Demand Response - Average Peak Load Reduction . uuycEyaycI,ycyyceycayco ): This KPI . n kW) measures the average reduction in total HVAC power consumption during active DR event periods compared to a defined baseline power consumption level that would have occurred without DR intervention . verage power during similar non-DR peak hours or a simulated "no-DR" Demand Response - Total Energy Saved/Shifted ( yuuycEyaycI,ycyyceycayco )): This metric . n kW. quantifies the total net reduction or shifting of energy consumption achieved specifically during all DR event periods throughout the simulation month, calculated by comparing the actual energy consumed during DR events with the energy that would have been consumed under a non-DR operational baseline. GP Policy Complexity: yuuyayaycI To assess the interpretability and conciseness of the evolved solutions, the complexity of the final selected GP policy is quantified by the total number of nodes . unctions and terminal. in its constituent expression tree. Results and discussion This section presents the empirical findings from the application of the proposed Genetic Programming (GP) framework and its comprehensive comparative evaluation against established baseline controllers. The analysis begins with an examination of the GP evolutionary process and the characteristics of the resultant policies, followed by a detailed benchmarking of performance across key metrics including energy efficiency, thermal comfort, and Demand Response (DR) effectiveness, and concludes with a statistical validation of the observed performance differentials. GP Evolutionary Process and Characteristics of Evolved Policies The foundation of the GP framework's success rests on the efficacy of its learning process. Before analyzing the final controller's performance, it is crucial to validate that the evolutionary algorithm effectively navigated the vast search space to discover and refine superior control policies. This ensures the final result is the product of a robust optimization process, not a random outcome. Fig. 9 provides the visual evidence of this successful learning journey over 100 generations. A detailed analysis of Fig. tracks the convergence of the three primary objective components for the best-performing individual in each generation, and All objectives, which are formulated for minimization, exhibit a clear and significant improvement. The initial, randomly generated policies perform poorly, as shown by the high penalty scores at generation 0. However, substantial reductions in both the . reen lin. range lin. penalties are observed within the initial 50 generations. This signifies that the best individuals progressively learned to achieve better thermal comfort and more effective Demand Response Notably, the . rimson lin. penalty stabilizes at a minimized value relatively early in the process. This suggests the GP algorithm quickly identified a baseline for energy- ISSN: 2252-4940/A 2025. The Author. Published by CBIORE S. Waheed and S. Int. Renew. Energy Dev 2025, 14. , 1221-1234 | 1229 Fig. 9 Programming Evolutionary Process and Convergence: . Convergence of the individual fitness objective components for the best . Convergence of the overall fitness score for both the best individual and the population average. efficient operation and then dedicated the majority of its subsequent evolutionary effort to fine-tuning the more complex and dynamic trade-offs between occupant comfort and grid Further illustrating the overall optimization progress. Fig. shows the performance of both the elite individuals and the population as a whole. A consistent downward trajectory is observed for both the best and average overall fitness scores, particularly within the initial 40-60 generations. This confirms successful population-wide learning and optimization towards better aggregate solutions. The narrowing gap between the two lines is particularly important, as it indicates that beneficial genetic material was being successfully distributed throughout the population, raising the quality of the entire gene pool. This robust, dual-level convergence at both the individual objective level and the aggregate population level validates the GP framework as a potent and reliable methodology for discovering holistically optimized solutions. Comparative Performance of Energy. Comfort, and GridResponsiveness Having established the robust convergence of the evolutionary process in the preceding section, the analysis now shifts to the performance of the final, representative GP policy selected from the Pareto front. This controller was rigorously benchmarked against the state-of-the-art DRL agent and the ASHRAE standards under identical operational conditions for the entire simulated month of July. Table 1 provides a comprehensive summary of the key performance indicators (KPI. across all three primary objectives energy efficiency, thermal comfort, and Demand Response effectiveness. A primary benchmark for evaluation is energy performance, where a distinct hierarchy between the controllers is immediately apparent from the data in Table 1. The GP policyAos total consumption of 6,800 kWh represents a substantial 9% reduction over the A2006 baseline and a 28. 4% saving over the G36 standard. Most critically, the 9. 3% energy saving relative to the state-of-the-art DRL (SAC) baseline highlights its superior energy optimization capabilities. Fig. 10a provides a clear, month-long visualization of this hierarchy, showing the GP controller's cumulative energy use consistently tracking below all others. The hourly operational dynamics depicted in the Fig. 10b heatmaps offer further insights into how these savings were achieved. Both the GP and DRL controllers effectively curtailed energy use during shoulder periods . :00-09:00 and 16:00-18:. , as indicated by the predominantly darker, lowerenergy regions. This suggests a more adept part-load operation Table 2 Statistical Significance of Key Performance Indicator (KPI) Differences Between Controllers . -value. Controller Pair 1 Controller Pair 2 Statistical Test UsedA p-value Total HVAC Energy Consumption . Evolved GP Policy ASHRAE A2006 Paired t-test < 0. Evolved GP Policy ASHRAE G36 Paired t-test Evolved GP Policy DRL (SAC) Paired t-test DRL (SAC) ASHRAE G36 Paired t-test ZAT Violation Degree-Hours (ACAh. Evolved GP Policy ASHRAE A2006 Paired t-test < 0. Evolved GP Policy ASHRAE G36 Paired t-test Evolved GP Policy DRL (SAC) Paired t-test DRL (SAC) ASHRAE G36 Paired t-test DR Peak Load Reduction . W) Evolved GP Policy ASHRAE G36* Independent t-test < 0. Evolved GP Policy DRL (SAC) Paired t-test DRL (SAC) ASHRAE G36* Independent t-test < 0. ISSN: 2252-4940/A 2025. The Author. Published by CBIORE Significance LevelA *** *** *** *** Waheed and S. Int. Renew. Energy Dev 2025, 14. , 1221-1234 | 1230 Fig. 10 Comparative Performance of Control Strategies: . Cumulative monthly HVAC energy consumption. Hourly energy consumption . Zone air temperature control on representative days. System dynamic response during a peak Demand Response event. and aggressive utilization of energy-saving measures. In stark contrast, the ASHRAE A2006 panel shows consistently high energy consumption across the entire occupied daytime block, indicative of its less adaptive, more rigid control logic. Crucially, these substantial energy savings were achieved without compromising occupant comfort. As the data in Table 1 confirms, the Evolved GP Policy achieved a notably low ZAT Violation Degree-Hour value of 75 ACAhr and maintained zone temperatures within the desired . AC-25AC] comfort band for 8% of all occupied hours. This high level of thermal comfort was statistically comparable to the DRL baseline . ACAhr, 5%) and significantly better than the ASHRAE G36 . ACAh. and A2006 . ACAh. The dynamic thermal performance is further elucidated in Fig. 10c, which presents indoor air temperature profiles for representative summer days. It is visually evident that the GP and DRL policies consistently regulate zone temperatures more tightly within the comfort band, exhibiting smoother profiles with minimal overshoot or In contrast, the ASHRAE baselines display more pronounced temperature fluctuations and larger deviations outside the designated comfort band. This visual evidence strongly suggests that the learning-based approaches achieve superior comfort stability due to their ability to learn more nuanced and anticipatory responses to varying load conditions. Finally, the GP policy's capacity to actively participate in Demand Response (DR) events was exceptional. summarized in Table 1, the controller achieved an average peak load reduction of 13. 7 kW . 1%), a performance comparable to the DRL baseline while the standard ASHRAE baselines showed negligible active participation. The dynamic response to a representative DR event is visualized in Fig. 10d, providing a clear, step-by-step illustration of the learned strategy. Upon activation of the DR signal, both intelligent controllers aggressively increased their Supply Air Temperature (SAT) setpoints from approximately 12. 5AC to 15. 5AC. As a direct result, total HVAC power consumption decreased dramatically from a peak of 17-19 kW to a minimal 3-4 kW for the duration of the event. During this period, the average zone temperatures experienced a controlled, slow drift, peaking around 25. 8AC before being brought back down. This clearly illustrates the GP framework's ability to evolve effective, explicit strategies for demand-side management, successfully completing the trifecta of holistic, multi-objective optimization. ISSN: 2252-4940/A 2025. The Author. Published by CBIORE S. Waheed and S. Int. Renew. Energy Dev 2025, 14. , 1221-1234 | 1231 Analysis of Evolved GP Operational Strategies. Policy Complexity, and Interpretability The superior performance documented in the previous section is not a mystery. A unique and powerful advantage of the Genetic Programming approach is the inherent transparency of its resulting policies. While the DRL agent's logic remains an opaque "black-box," the fundamental structure of the evolved GP policies, represented as expression trees, allows for direct inspection, analysis, and human understanding. This section deconstructs the evolved strategies to reveal the drivers of its success and discuss the critical importance of First, it is important to address the complexity of the evolved solution. The complete multi-output GP policy, comprising distinct expression trees for each of the four AHU control variables, aggregates to a total of 185 nodes. This moderate complexity represents a successful balance: the policy is sophisticated enough to execute nuanced, high-performance control, yet it remains entirely human-inspectable. This indicates that the GP evolved functionally effective policies without an unmanageable degree of structural complexity, thereby preserving interpretability. Fig. 11 offers a multi-faceted visualization of the controller's evolved "intelligence," revealing the specific, learned strategies that led to its superior performance. Fig. visualizes the multi-dimensional control strategy for the AHU SAT setpoint. The surface reveals a distinctly non-linear and adaptive strategy. When zones are at or below their setpoint . ero or negative erro. , the GP maintains a higher, energysaving SAT . pproximately 15-16. 5AC). However, as zones become warmer, the GP policy enacts a sharp reduction in the SAT, driving it towards its lower operational limit of approximately 11AC when the error reaches 2. 0AC. This aggressive cooling response demonstrates a learned prioritization, a sophisticated, threshold-influenced behavior that would be challenging to hand-craft. Further insights into emergent operational behavior are provided in Fig. 11b and 11c clearly delineates distinct operational modes: a dense cluster of points shows operation with a low SAT . 5AC-12. 5AC) and an Fig. 11 Analysis of Evolved GP Operational Strategies: . Control surface for the AHU SAT setpoint. Emergent coordination of AHU and chiller plant operation. Intelligent economizer free cooling logic. Daily activation patterns of Supply Water Temperature (SWT) management regimes. ISSN: 2252-4940/A 2025. The Author. Published by CBIORE S. Waheed and S. Int. Renew. Energy Dev 2025, 14. , 1221-1234 | 1232 aggressive, low Chiller SWT . -9AC, reddish hue. , indicating active cooling under high demand. A sparser cluster at higher SATs . -16AC) corresponds to a higher, more efficient Chiller SWT . -12AC, bluish hue. , representing an energy-saving mode during lower loads. This demonstrates a learned coordination between the AHU and chiller plant. A high outdoor air fraction . -100%) is predominantly utilized when the outdoor air is cool . -20AC). As the outdoor temperature surpasses 22AC, the GP decisively minimizes the outdoor air fraction to avoid introducing excessive thermal loads. Finally, to understand the temporal dynamics. Fig. 11d presents a heatmap illustrating which specific "GP Rule/Regime for SWT" is active at each timestep. The color coding clearly distinguishes these For example, on high-load days like 2023-07-04, a clear shift from inactive . egime '0') to active cooling . egimes '4', '5', '6', indicated by yellow/orange/gree. is evident during the morning ramp-up and peak afternoon hours. Conversely, during milder periods or weekends, lower-numbered regimes ('0', '1', '2', white/pin. , indicating reduced cooling, are predominantly active. This visualization highlights the GP's ability to dynamically switch between operational modes based on the time of day and prevailing conditions, providing a granular view of its adaptive behavior beyond static setpoint This ability to deconstruct and verify the control logic is the paramount contribution of this work. It directly addresses the primary barrier to the adoption of advanced AI in critical building systems: a lack of trust and verifiability (Yu et al. , 2. While a "black-box" model may perform well, its opacity creates significant implementation hurdles. The GP controller's "whitebox" nature provides the transparency necessary for vetting, trust, and practical implementation by facility managers, offering a viable pathway to deploy truly intelligent and trustworthy building automation. Statistical Significance of Performance Differences To ascertain the statistical robustness of the observed performance advantages, a comprehensive analysis was conducted using 10 independent simulation runs for each controller, with the results summarized in Table 2. This analysis formally validates that the superior performance of the evolved GP policy is not an artifact of a single simulation but a consistent and statistically significant outcome. In the critical domain of energy efficiency, the analysis reveals a clear hierarchy. The GP controller was found to be significantly superior to all other methods, consuming less energy than the ASHRAE A2006 . < . and G36 . = 0. Most consequentially, the 3% energy saving achieved by the GP policy over the state-ofthe-art DRL agent was also confirmed to be a statistically significant advantage . = 0. This finding is pivotal, as it provides strong empirical evidence that, for this complex control problem, the GP's evolutionary search discovered a more globally optimal policy than the DRL's gradient-based While achieving this superior efficiency, the GP policyAos performance in maintaining thermal comfort and executing demand response was statistically indistinguishable from the DRL agent, with p-values of 0. 350 and 0. This demonstrates that the GPAos energy advantage was not a simple trade-off but a genuine optimization gain, achieved without any statistically significant sacrifice in other key performance areas. These statistical findings are highly significant because they paint a clear, empirically-backed picture of a holistically superior and more intelligent controller. The GP model did not simply find a good solution. it found a better way to balance the system's competing objectives. The fact that its energy savings are statistically significant while its comfort and DR performance remain on par with the DRL agent suggests that the GP learned to successfully decouple energy efficiency from comfort provision in a way the DRL agent could not. The true importance of our model, however, is realized when coupling this statistically validated performance with its inherent interpretability, as discussed in the previous section. This combination directly addresses the primary barrier to the adoption of advanced AI in critical building systems: the blackbox problem, which creates a fundamental lack of trust and verifiability (Cpalka. Aapa & PrzybyC, 2. By providing a solution that is not only statistically proven to be more efficient but is also fully transparent and auditable, our GP model represents a viable and highly advantageous pathway toward intelligent HVAC control that is not just effective, but also trustworthy and practically deployable in real-world HVAC control. Limitations and Future work While this study robustly demonstrates the significant potential of the GP-evolved controllers, several considerations for practical application and future research warrant discussion. The findings, derived from a high-fidelity simulation environment, provide a strong performance benchmark, but onsite validation is the logical next step to confirm performance against real-world dynamics. A key methodological limitation is the lack of a separate validation dataset, which raises the potential for overfitting. Future work must therefore validate the evolved policies against different weather years and seasons to rigorously assess their generalization performance and Translating this research into industry practice requires addressing two primary barriers: the integration with proprietary Building Automation Systems (BAS) via standardized APIs, and the upfront computational and expertise requirements for policy evolution. However, the inherent transparency of GP offers a clear advantage over opaque AI, significantly lowering these adoption hurdles. A practical pathway could involve a Control-as-a-Service model, where foundational policies are evolved for building archetypes and then presented to facility managers for verification. This whitebox nature enables a phased, trust-building deployment like shadow mode with human oversight, contrasting sharply with the all-or-nothing trust demanded by black-box DRL agents. Looking ahead, future work should explore hybrid models that combine the strengths of GP and Deep Reinforcement Learning. A promising approach involves using GP to evolve an interpretable, high-level strategic frameworkAidefining the operational modes and primary logic while a DRL agent is tasked with fine-tuning the continuous control parameters within that GP-defined logic in real-time. Such a hybrid system could offer the best of both worlds: the robust, transparent, and verifiable strategic intelligence of GP, coupled with the adaptive, fine-grained optimization of DRL. Conclusion This research successfully pioneered and rigorously validated a Genetic Programming (GP) framework for the direct evolution of interpretable, multi-objective HVAC control policies, addressing the critical black-box limitations of contemporary AI controllers while integrating sophisticated Demand Response ISSN: 2252-4940/A 2025. The Author. Published by CBIORE S. Waheed and S. Int. Renew. Energy Dev 2025, 14. , 1221-1234 | 1233 Comprehensive simulations within a validated multi-zone office building model demonstrated that the GPevolved policies achieved superior energy efficiency, reducing total HVAC consumption by a significant 40. 9% against ASHRAE A2006, 28. 4% over ASHRAE G36, and a 3% compared state-of-the-art Deep Reinforcement Learning agent, all while maintaining excellent thermal comfort for 98. 8% of occupied hoursAia level comparable to DRL and markedly better than standard Furthermore, the GP policies exhibited robust DR effectiveness, delivering a 72. 1% peak load reduction through learned strategies like pre-cooling and dynamic setpoint modulation, performing on par with DRL. The paramount contribution of this work is the attainment of this high operational performance through policies that are inherently transparent 185 total nodes for the complete AHU strategy, allowing for direct human inspection, verification, and trust. This contrasts sharply with opaque DRL models and obviates the need for potentially inexact post-hoc explanations. directly evolving understandable, high-performing solutions. GP is substantiated as a potent and practical methodology for advancing intelligent building automation towards more efficient, grid-responsive, and trustworthy systems. Future investigations should prioritize real-world deployment, enhancing GP scalability for larger systems, and evolving adaptive policies with greater resilience to operational uncertainties and faults. References