A surrogate-assisted reinforcement learning framework for thermal comfort regulation in subway cabins
Journal article, 2026

Intelligent thermal comfort regulation in enclosed transit cabins requires efficient decision-making under coupled airflow, heat transfer, occupant thermal response, and operational constraints. Computational fluid dynamics (CFD)-based evaluation provides detailed physical information, but its computational cost limits direct use in repeated regulation-oriented evaluations. Addressing this issue, this paper proposes a surrogate-assisted reinforcement learning framework for sequential thermal comfort regulation of an impinging jet ventilation (IJV) system in a subway cabin. A validated CFD model, calibrated against full-scale measurements in an actual subway cabin, was first used to generate offline data under different operating conditions. Based on these data, six Multi-Layer Perceptron (MLP) surrogate models were trained to predict passenger-wise thermal discomfort at six representative locations. The regulation problem was then formulated as a Markov decision process, and a Deep Q-Network (DQN) was trained to sequentially adjust the controllable supply parameters for regulation. Numerical experiment results demonstrated that all surrogate models provided usable predictive accuracy across all monitored positions, and the learned DQN policy achieved a 97% success rate, reduced the average maximum discomfort from 30.50% to 12.51%, and required 25 steps on average to reach the target region. Further CFD validation under four representative regulated operating conditions showed good agreement with the framework predictions, with a maximum average deviation of only 0.673% between predicted values and CFD results. Despite inherent local variations induced by complex jet entrainment, the intelligent regulation effectively maintains macroscopic spatial uniformity, strictly ensuring the global consistency of occupant thermal comfort. By dynamically shifting regulation strategies, the algorithm achieves an optimal trade-off. It leverages low-volume modes for deep cooling, while utilizing high-volume modes to mitigate vertical thermal stratification via enhanced turbulent mixing. Ultimately, this dynamic shifting ensures a robust balance between maximizing convective heat dissipation and preventing localized draft risks, establishing a highly versatile, mechanism-driven intelligent paradigm for enclosed space thermal management. For other enclosed environments, the workflow can be re-instantiated by redefining controllable variables, monitored comfort indicators, operating bounds, and reward constraints.

Surrogate model

Impinging jet ventilation

Subway thermal comfort management

Reinforcement learning

Fiala thermoregulation model

Author

Puyang Zhang

Central South University

City University of Hong Kong

Zhexi Yang

City University of Hong Kong

Xifeng Liang

Central South University

Xinchao Su

Central South University

Guangjun Gao

Central South University

Sinisa Krajnovic

Chalmers, Mechanics and Maritime Sciences (M2), Fluid Dynamics

Advanced Engineering Informatics

1474-0346 (ISSN)

Vol. 76 105078

Subject Categories (SSIF 2025)

Computer Sciences

Energy Engineering

Applied Mechanics

DOI

10.1016/j.aei.2026.105078

More information

Latest update

7/29/2026