Transferable reinforcement learning for demand-responsive transit scheduling via offline pre-training and online fine-tuning
Journal article, 2027

Demand-responsive transit (DRT) offers flexibility in adapting to dynamic travel demands and improving resource allocation efficiency. However, existing scheduling models suffer from limited generalization and the cold-start problem when deployed in new operational contexts. This study adapts reinforcement learning to cross-region DRT scheduling by combining offline imitation learning with online proximal policy optimization in a transferable framework (TR-PPO). The proposed approach follows a two-stage learning paradigm. In the first stage, behavior cloning is conducted on extensive expert datasets to learn transferable dispatching patterns and establish an initial policy prior. By using relative spatial representations rather than absolute coordinates, the framework reduces dependence on absolute coordinates and improves cross-region feature transfer. In the second stage, the pre-trained model is fine-tuned online in the target region using PPO to improve policy adaptation to local spatiotemporal dynamics under limited-data conditions. Case study findings from real-world datasets in China demonstrate that TR-PPO outperforms the tested scheduling baselines, achieving total system cost reductions of 11.6%–27.7%. Across the tested within-city, single-source cross-city, and multi-source cross-city transfer settings, the transfer-based policies also achieve competitive performance relative to the other baseline algorithms. TR-PPO accelerates convergence during online fine-tuning, reducing the number of training episodes by 57.8%. Moreover, TR-PPO demonstrates strong data efficiency during online fine-tuning, achieving competitive performance even with only 25% of the target-region training data. The results indicate that the proposed framework can improve the adaptation efficiency of DRT scheduling policies across the tested source-target regions with different demand patterns and network structures.

Online fine-tuning

Reinforcement learning

Offline pre-training

Bus scheduling

Demand-responsive transit

Author

Jing Bian

Beihang University

Xiaolei Ma

Beihang University

Haoyang Yan

Beihang University

Hesham El-Sayed

United Arab Emirates University

Kun Gao

Chalmers, Architecture and Civil Engineering, Geology and Geotechnics

Xiaohan Liu

Chalmers, Architecture and Civil Engineering, Geology and Geotechnics

Transportation Research, Part C: Emerging Technologies

0968-090X (ISSN)

Vol. 194 106010

Subject Categories (SSIF 2025)

Transport Systems and Logistics

Computer Sciences

DOI

10.1016/j.trc.2026.106010

More information

Latest update

9/22/2026