Deploying hundreds of robots in a warehouse, where conditions change by the minute, is a complex task. Moreover, exact solutions such as the Hungarian algorithm are too slow to be used in real-time. In classical Reinforcement Learning (RL), the above-mentioned state-of-the-art methods hardly scale to learning in very high-dimensional spaces.
We name our method as GNN-IL-PPO, which is a composite policy by GNN and imitation pre-training, together with PPO fine-tuning. We describe the warehouse as a heterogeneous graph having three kinds of nodes, but our policy is scalable to more robots. The trained task-efficient-urgent bimodal heuristic expert is then bootstrapped with demonstrations, and the requirements of exploration are reduced. PPO further generalizes the policy to new tasks.
In simulations with 15 robots, the proposed method has comparable performance as state-of-the-art in terms of makespan loss (up to 8%) while running orders of magnitude faster. Ablations indicate the importance of the graph structure and pre-training, which are 2x faster converge than training from scratch.
Novel graph representation integrating robot states, tasks, and warehouse layout with typed nodes and edges for permutation-invariant, scalable learning.
Graph Attention Networks for local neighborhood aggregation combined with Transformer layers for global reasoning and coordination patterns.
Two-phase pipeline using behavior cloning from expert heuristics followed by PPO fine-tuning for sample-efficient learning with long-term optimization.
Comprehensive experiments comparing against state-of-the-art methods with detailed ablation studies validating each component's contribution.
| Method | Makespan (s) | Success (%) | Throughput | Time (s) |
|---|---|---|---|---|
| Hungarian (Optimal) | 58.2 ± 2.1 | 100.0 | 0.86 ± 0.03 | 1.82 |
| GNN-IL-PPO (Ours) | 62.8 ± 2.5 | 94.2 ± 1.8 | 0.80 ± 0.04 | 0.72 |
| GNN-PPO (Ablation) | 66.5 ± 3.2 | 88.5 ± 2.5 | 0.75 ± 0.05 | 0.71 |
| Greedy-Nearest | 68.5 ± 3.8 | 78.5 ± 3.2 | 0.73 ± 0.05 | 0.18 |
| Random | 85.2 ± 5.5 | 62.5 ± 4.5 | 0.59 ± 0.06 | 0.05 |