交通运输系统工程与信息 ›› 2026, Vol. 26 ›› Issue (4): 380-390.DOI: 10.16097/j.cnki.1009-6744.2026.04.033

• 智慧机场运营管理 • 上一篇    

基于深度强化学习的机场地面保障调度方法

程华1,张扬*1,向飞2,马吉震2,唐明杰2   

  1. 1. 中国民用航空总局第二研究所,成都 610042;2. 民航成都信息技术有限公司,成都 610042
  • 收稿日期:2025-12-25 修回日期:2026-02-06 接受日期:2026-03-12 出版日期:2026-08-25 发布日期:2026-08-21
  • 作者简介:程华(1973— ),男,四川成都人,研究员
  • 基金资助:
    民航安全能力建设中试项目 (RJ2025076)

Airport Ground Handling Scheduling Approach Based on Deep Reinforcement Learning

CHENG Hua1, ZHANG Yang*1, XIANG Fei2, MA Jizhen2, TANG Mingjie2   

  1. 1. The Second Research Institute of Civil Aviation Administration of China, Chengdu 610042, China; 2. Civil Aviation Information Technology Co Ltd of Chengdu, Chengdu 610042, China
  • Received:2025-12-25 Revised:2026-02-06 Accepted:2026-03-12 Online:2026-08-25 Published:2026-08-21
  • Supported by:
    Pilot Project for Civil Aviation Safety Capacity Building (RJ2025076)

摘要: 地面保障涵盖航班落地至起飞的关键作业环节,各环节任务间存在条件时序依赖,同时,多航班作业过程伴随资源竞争,这种复杂耦合关系成为制约地面保障发展的关键难题。传统方法依赖人工规则设计,在大规模和高动态场景下难以适用。本文提出一种融合数学图嵌入与注意力机制的深度强化学习算法框架,将具有随机任务序列的动态调度问题通过条件任务图形式化建模为混合整数规划模型,并将其转化为异质图,利用图卷积网络提取全局图嵌入以高效表征图中的耦合关系,与实时状态融合形成深度强化学习状态表示,通过基于注意力机制的策略网络学习序列化调度策略,实现端到端调度策略优化。在真实机场仿真环境中的测试结果表明:该算法框架与传统启发式方法以及强化学习方法相比,在多个核心指标上均有显著提升,其中,与深度Q网络相比,航班平均延误降低40.1%,资源利用率提升6.6%,验证了算法的有效性与优越性。

关键词: 航空运输, 机场地面保障, 深度强化学习, 保障调度, 条件任务图, 混合整数规划模型, 图卷积网络

Abstract: Airport ground handling covers the key operational links from flight landing to takeoff, and there are conditional temporal dependencies between tasks in each link. At the same time, the multi flight operation process is accompanied by resource competition, and this complex coupling relationship has become a major challenge restricting the development of ground support. Traditional methods rely on manual rule design and are difficult to apply in large- scale and high dynamic scenarios. This paper proposes a deep reinforcement learning algorithm framework that integrates mathematical graph embedding and attention mechanism for the airport ground handling scheduling. The dynamic scheduling problem with random task sequences is graphically modeled as a mixed integer programming model through conditional task graphs, and transformed into a heterogeneous graph. The graph convolutional network is used to extract global graph embedding to efficiently represent the coupling relationship in the graph, which is fused with real-time state to form a deep reinforcement learning state representation. The attention mechanism based policy network is used to learn serialized scheduling strategies, achieving end-to-end scheduling strategy optimization. The test results in a real airport simulation environment show that compared with traditional heuristic methods and reinforcement learning methods, this algorithm framework has significant improvements in multiple core metrics. Among them, compared with deep Q-networks, the average flight delay is reduced by 40.1%, and the resource utilization rate is improved by 6.6%, which verifies the effectiveness and superiority of the algorithm.

Key words: air transportation, airport ground handling, deep reinforcement learning, handling scheduling, conditional task graph; mixed integer programming model, graph convolutional network

中图分类号: