A Spatio-temporal Transformer-STGNN Hybrid Reinforcement Learning Method for Bus Cooperative Optimization
-
摘要: 为解决公交动态越站与驻站多策略协同中局部决策引发全局连锁延误及多重目标冲突的难题,研究了1种融合Transformer与时空图神经网络的动态协同优化模型。构建双向交互的感知与预知时空特征提取架构。利用时空变换器的空间注意力机制提取公交站点的动态依赖关系,生成自适应权重矩阵并作为动态邻接矩阵输入时空图神经网络。该网络结合图卷积网络与因果扩张卷积,预测客流积压与延误传播风险,并将预知风险评分反向传递至时序注意力层。此闭环反馈机制通过风险驱动调整权重,量化了单一调度决策产生的时空非线性连锁效应。设计了1种嵌套改进遗传算法与深度双Q网络的混合强化学习框架。利用遗传算法的全局广度搜索生成帕累托前沿解集,将其作为深度双Q网络的初始策略空间;同时,在多目标奖励函数中引入时空图神经网络的风险评分作为安全约束惩罚,以此权衡车辆运行效率与乘客出行成本,输出动态越站与驻站协同控制策略。基于佛山市101路公交线路高峰时段运营数据开展仿真实验。结果表明,该策略控制了车头时距波动,避免了公交车辆的连续串车现象。在乘客平均等待时间增加0.83%的前提下,乘客平均在途时间减少约24.7%,同时缩短了车辆整体在途时间并提升了站间行驶速度。相比传统单一深度强化学习算法,该混合算法具备更少的迭代次数与更小的寻优误差。适用于高频发车的城市干线公交实时协同调度,能在控制时空风险传播的前提下实现多调度策略的动态平衡。Abstract: To address multi-objective conflicts and global ripple delays caused by local decisions, a dynamic cooperative optimization model is investigated. This model integrates a Spatio-Temporal Transformer and a spatiotemporal graph neural network (STGNN). A bidirectional interactive architecture for the perception and anticipation of spatiotemporal features is constructed. Dynamic dependencies of bus stops are extracted using the spatial attention mechanism of the Transformer. An adaptive weight matrix is generated and subsequently inputted into the STGNN as a dynamic adjacency matrix. This network combines graph convolutional networks and causal dilated convolutions to predict risks of passenger accumulation and delay propagation. The anticipated risk scores are reversely transmitted to the temporal attention layer. Driven by risks, this closed-loop feedback mechanism adjusts the weights. Thereby, the nonlinear ripple effects in space and time generated by single scheduling decisions are quantified. A hybrid reinforcement learning framework is designed. An improved genetic algorithm (GA) and a Deep Double Q-Network (DDQN) are nested in this framework. A Pareto front solution set is generated using the global broad search of the GA. This set serves as the initial strategy space for the DDQN. Simultaneously, the risk scores from the STGNN are introduced into the multi-objective reward function as safety constraint penalties. Consequently, the operational efficiency of vehicles and the travel costs of passengers are balanced. Thus, optimal coordinated control strategies for stop-skipping and holding are generated. Simulation experiments are conducted based on the operational data of Foshan Bus Route 101 during peak hours. The results indicate that the fluctuations in headways are controlled by this strategy. Furthermore, the continuous bus bunching phenomenon is successfully avoided. Under the premise of a 0.83% increase in the average passenger waiting time, the average passenger in-vehicle time is reduced by approximately 24.7%. Meanwhile, the overall vehicle travel time is shortened, and the driving speed between stops is improved. Compared with traditional single deep reinforcement learning algorithms, this hybrid algorithm exhibits fewer iteration counts and smaller optimization errors. The proposed model is applicable to the real-time cooperative scheduling for urban trunk buses with high-frequency departures. A dynamic balance of multiple scheduling strategies is achieved under the premise of controlling risk propagation.
-
表 1 实验场景设计说明
Table 1. Design of experimental scenarios
场景 场景编号 发车间隔/min 协同优化 GA GA-DDQN 高峰时段 1 8 × × × 2 8 √ × × 3 8 √ √ × 4 8 √ × √ 表 2 不同实验公交在途时间分析表
Table 2. Analysis table of in-transit time of different experimental buses
场景 总在途时间/s 平均在途时间/s 高于平均时间车辆数 无协同 36 364.5 4 010.5 5 仅协同 36 037.8 4 004.2 5 采用GA 34 295.4 3 810.6 5 采用GA-DDQN 34 090.2 3 787.8 4 表 3 不同实验公交行驶速度分析表
Table 3. Analysis table of different experimental bus speeds
场景 站间速度中位数/(km/h) 平均速度/(km/h) 无协同 17.1 10.76 仅协同 17.4 10.83 采用GA 18.3 10.92 采用GA-DDQN 19.1 11.41 -
[1] 彭刚. 基于多源数据的公交客流预测及发车时刻表优化方法研究[D]. 大连: 大连海事大学, 2023.PENG G. Research on bus passenger flow prediction and departure timetable optimization method based on multi-source data[D]. Dalian: Dalian Maritime University, 2023. (in Chinese) [2] 焦峰. 基于多源数据融合的城市公交客流预测与预警研究[D]. 北京: 北京交通大学, 2023.JIAO F. Research on urban bus passenger flow prediction and early warning based on multi-source data fusion[D]. Beijing: Beijing Jiaotong University, 2023. (in Chinese) [3] 孙世超, 吕豪. 大数据环境下基于职住地识别的公交通勤行为判断与特征分析[J]. 上海海事大学学报, 2023, 44(4): 45-50.SUN S C, LYU H. Public transport commuting behavior judgment and characteristic analysis based on workplace and residence identification under big data environment[J]. Journal of Shanghai Maritime University, 2023, 44(4): 45-50. (in Chinese) [4] 邓红星, 韩树鑫. 基于手机信令数据的城市居民出行特征季节性差异分析[J]. 重庆理工大学学报(自然科学), 2022, 36 (9): 202-210.DENG H X, HAN S X. Seasonal difference analysis of urban residents' travel behavior characteristics based on mobile signaling data[J]. Journal of Chongqing University of Technology (Natural Science), 2022, 36(9): 202-210. (in Chinese) [5] 王鹏飞. 实时供需信息条件下城市公交运营优化策略研究[D]. 南京: 东南大学, 2022.WANG P F. Research on optimization strategies for urban bus operations under real-time supply-demand information[D]. Nanjing: Southeast University, 2022. (in Chinese) [6] 赵小梅, 朱香原, 汪秦, 等. 多类型公交线路的运行可靠性优化策略[J]. 华南理工大学学报(自然科学版), 2023, 51(8): 32-39, 50.ZHAO X M, ZHU X Y, WANG Q, et al. Operational reliability optimization strategies of multi-type bus lines[J]. Journal of South China University of Technology (Natural Science Edition), 2023, 51(8): 32-39, 50. (in Chinese) [7] 卢凯, 夏小龙, 胡建伟, 等. 基于交叉口相位裕量时间的公交准点控制模型[J]. 同济大学学报(自然科学版), 2019, 47 (12): 1742-1747.LU K, XIA X L, HU J W, et al. Bus punctuality control model based on the phase margin time at intersections[J]. Journal of Tongji University (Natural Science), 2019, 47(12): 1742-1747. (in Chinese) [8] 温惠英, 张东冉, 陆思园. GA-LSTM模型在高速公路交通流预测中的应用[J]. 哈尔滨工业大学学报, 2019, 51(9): 81-87, 95.WEN H Y, ZHANG D R, LU S Y. Application of GA-LSTM model in highway traffic flow prediction[J]. Journal of Harbin Institute of Technology, 2019, 51(9): 81-87, 95. (in Chinese) [9] CHAO P, XU Y, HUA W, et al. A survey on map-matching algorithms[C]. Databases Theory and Applications: 31st Australasian Database Conference, Melbourne: Springer International Publishing, 2020. [10] CUI G, BIAN W, WANG X. Hidden Markov map matching based on trajectory segmentation with heading homogeneity[J]. GeoInformatica, 2021, 25: 179-206. doi: 10.1007/s10707-020-00429-4 [11] QIU Y, WANG J, JIN Z, et al. Pose-guided matching based on deep learning for assessing quality of action on rehabilitation training[J]. Biomedical Signal Processing and Control, 2022, 72: 103323. doi: 10.1016/j.bspc.2021.103323 [12] 周文竹, 汪琦, 王楠. 交通视角下防御单元的适应性规划策略——基于突发公共卫生安全事件的思考[J]. 城市交通, 2021, 19(6): 71-80, 90.ZHOU W Z, WANG Q, WANG N. adaptive planning strategy of defense unit from traffic perspective: thinking based on public health emergencies[J]. Urban Transport of China, 2021, 19(6): 71-80, 90. (in Chinese) [13] 宋钰. 基于多源数据的城市交通需求时空格局演化及驱动机理[D]. 大连: 大连海事大学, 2023.SONG Y. Research on spatiotemporal pattern evolution and driving mechanisms of urban transportation demand using multi-source data[D]. Dalian: Dalian Maritime University, 2023. (in Chinese) [14] 卢凯, 田鑫, 林观荣, 等. 交叉口信号相位设置与配时同步优化模型[J]. 浙江大学学报(工学版), 2020, 54(5): 921-930.LU K, TIAN X, LIN G R, et al. Synchronous optimization model of intersection signal phase design and timing[J]. Journal of Zhejiang University (Engineering Science), 2020, 54 (5): 921-930. (in Chinese) [15] PRASSAS E S, ROESS P R. The highway capacity manual: a conceptual and research history volume 2[M]. Cham: Springer International Publishing, 2020. [16] LIU Z, YAN Y, QU X, et al. Bus stop-skipping scheme with random travel time[J]. Transportation Research Part C: Emerging Technologies, 2013, 35: 46-56. doi: 10.1016/j.trc.2013.06.004 [17] YU Y B, YE Z R, WANG C. Study of bus stop skipping scheme based on modified cellular genetic algorithm[M]. Reston: American Society of Civil Engineers, 2015. [18] CHEN X, HELLINGA B, CHANG C, et al. Optimization of headways with stop-skipping control: a case study of bus rapid transit system[J]. Journal of Advanced Transportation, 2015, 49(3): 385-401. doi: 10.1002/atr.1278 [19] MOU Z, ZHANG H, LIANG S. Reliability optimization model of stop-skipping bus operation with capacity constraints[J]. Journal of Advanced Transportation, 2020, 2020: 1-11. [20] ZHANG L, HUANG J, LIU Z, et al. An agent-based model for real-time bus stop-skipping and holding schemes[J]. Transportmetrica A: Transport Science, 2021, 17(4): 615-647. doi: 10.1080/23249935.2020.1802363 [21] GKIOTSALITIS K. Robust stop-skipping at the tactical planning stage with evolutionary optimization[J]. Transportation Research Record, 2019, 2673(3): 611-623. doi: 10.1177/0361198119834549 [22] HUANG D, XING J, LIU Z, et al. A multi-stage stochastic optimization approach to the stop-skipping and bus lane reservation schemes[J]. Transportmetrica A: Transport Science, 2021, 17(4): 1272-1304. doi: 10.1080/23249935.2020.1858206 [23] HE S X, LIANG S D, DONG J, et al. A holding strategy to resist bus bunching with dynamic target headway[J]. Computers & Industrial Engineering,2020, 140: 106237. [24] WANG J, SUN L. Dynamic holding control to avoid bus bunching: a multi-agent deep reinforcement learning framework[J]. Transportation Research Part C: Emerging Technologies, 2020, 116: 102661. doi: 10.1016/j.trc.2020.102661 [25] BIE Y, XIONG X, YAN Y, et al. Dynamic headway control for high-frequency bus line based on speed guidance and intersection signal adjustment[J]. Computer-Aided Civil and Infrastructure Engineering, 2020, 35(1): 4-25. doi: 10.1111/mice.12446 [26] 奇格奇, 曹琳琪, 沈益达, 等. 基于GNSS轨迹数据的公交多能源供需网络调度优化模型[J]. 交通信息与安全, 2025, 43(3): 85-99.QI G Q, CAO L Q, SHEN Y D, et al. An optimization model for multi-energy supply and demand network scheduling of public transportation based on GNSS trajectory Data[J]. Journal of Transport Information and Safety, 2025, 43(3): 85-99. (in Chinese) [27] 郑乐, 高良鹏, 沈金星, 等. 可变线路式公交运行服务能力优化策略[J]. 哈尔滨工业大学学报, 2021, 53(9): 107-115.ZHEN L, GAO L P, SHEN J X, et al. Operational service capability optimization strategies for flex-route transit service[J]. Journal of Harbin Institute of Technology, 2021, 53 (9): 107-115. (in Chinese) [28] ZHANG Y B, ZHOU X M, FAN W D, et al. A multi-agent deep reinforcement learning framework for coordinated optimization of bus operations[J]. Computers & Industrial Engineering, 2026, 214: 111895. [29] ZHOU X M, ZHANG Y B, WEI G H, et al. Collaborative optimization method for intelligent and connected single-route bus operation[J]. Journal of Intelligent Transportation Systems, 2025, 31(3): 1-19. [30] RODRIGUEZ J, KOUTSOPOULOS H N, WANG S H, et al. Cooperative bus holding and stop-skipping: a deep reinforcement learning framework[J]. Transportation Research Part C: Emerging Technologies, 2023, 155: 104308. doi: 10.1016/j.trc.2023.104308 -
下载: