露天矿软基环境下多无人卡车协同安全规划方法

Safety-aware cooperative planning for multi-agent autonomous trucks on open-pit soft terrains

  • 摘要: 露天矿软基环境的复杂性给多智能体强化学习带来严峻的挑战。传统的纯几何避障策略忽略底层车地交互力学,而现有的多智能体强化学习算法通常采用软惩罚机制且因联合更新引发环境非平稳性,在面对软基路面时表现出较差的安全性与协同性能。为解决上述问题,本文提出一种融合物理先验的强化学习方法,能够学习满足严格物理安全约束且充分利用全局环境信息的多卡车协同规划策略。该方法将Bekker-Wong地面力学模型融入受限马尔可夫博弈底层结构,在策略寻优中考虑软基环境的承载极限与物理安全边界。采用集成多头自注意力和卷积神经网络的异构双流感知网络,确保智能体能够同时处理高维空间物理场以及来自多车交互图的变长时序拓扑信息。设计一种带有信赖域对偶投影与动态优先级的顺序策略更新算法,能有效降低多车并发引起的非平稳性与死锁现象。仿真实验表明露天矿多车协同场景中,本文所提出的方法在物理安全性、收敛速度及累计路面载荷分布均衡性上均优于现有的多智能体受限强化学习方案。

     

    Abstract: The complexity of soft-terrain environments in open-pit mines poses severe challenges to multi-agent reinforcement learning (MARL). Traditional geometric obstacle-avoidance strategies neglect the underlying wheel-terrain interaction mechanics. Meanwhile, existing MARL algorithms typically rely on soft penalty mechanisms and suffer from environment non-stationarity induced by joint policy updates, leading to poor safety and coordination performance on soft terrains. To address these issues, we propose a physics-informed reinforcement learning framework to learn cooperative truck-planning policies that strictly satisfy physical safety constraints while fully leveraging global environmental contexts. Specifically, the Bekker-Wong terramechanics model is embedded into the underlying formulation of a constrained Markov game, explicitly accounting for the bearing capacity limits and physical safety boundaries of the soft subgrade during policy optimization. Furthermore, a heterogeneous dual-stream perception architecture integrating multi-head self-attention and Convolutional Neural Networks (CNNs) is developed, enabling the agents to simultaneously process high-dimensional spatial physical fields and variable-length temporal-topological information from multi-truck interaction graphs. Additionally, we design a sequential policy-update scheme equipped with trust-region dual projection and dynamic priority assignment, which effectively mitigates the non-stationarity and deadlock phenomena caused by concurrent multi-agent updates. Simulation results in open-pit multi-truck coordination scenarios demonstrate that the proposed method consistently outperforms existing constrained MARL baselines in terms of physical safety, convergence speed, and spatial uniformity of cumulative traffic-load distribution.

     

/

返回文章
返回