论文部分内容阅读
为改善交叉口排队长度管理,避免交叉口某个方向排队长度过长,采用强化学习理论建立了以平均排队长度差最小为优化目标的在线Q学习模型。针对控制性能指标相对于邻近的配时方案不敏感的特点,提出了以平均排队长度差作为基本单位重新构造奖励函数,目的是拉大各行为对应的Q值差距,提高模型的收敛速度和鲁棒性。集成Excel VBA,Vissim,Matlab建立了在线仿真平台,作为计算环境对算例进行了计算。算例中利用GPS数据对Vissim软件中车辆加减速度曲线进行了标定。计算结果表明以平均排队长度差作为优化目标能够提高各个方向排队长度的平衡性,优化整个交叉口的时空资源;建立的在线Q模型具有学习能力和较快的计算速度,模型能否收敛受到周期取值和可选行为数量的影响。
In order to improve the queue length management at the intersection and avoid the queue length in one direction at the intersection, the online Q learning model with the minimum average queue length difference as the optimization objective was established by using the reinforcement learning theory. In view of the characteristics that the control performance index is insensitive to the adjacent timing scheme, this paper proposes to reconstruct the reward function with the average queuing length difference as the basic unit, in order to widen the Q value gap of each behavior and improve the convergence speed of the model Great. Integrated Excel VBA, Vissim, Matlab established an online simulation platform, as a computing environment for the calculation of the example. The example uses GPS data to calibrate the acceleration and deceleration of vehicles in Vissim software. The results show that using the average queue length difference as the optimization target can improve the balance of queue length in all directions and optimize the space-time resources of the whole intersection. The established online Q model has learning ability and fast calculation speed, The value and the number of optional behavior.