Abstract
Autonomous driving in complex urban environments requires trajectory planning that balances safety, efficiency, and human-like behavior. Although imitation learning (IL) can capture expert driving patterns from large-scale demonstrations, existing IL-based planners still face challenges in safety-critical scenarios and long-tail traffic distributions. Meanwhile, optimization-based planners provide explicit constraint handling but are often separated from upstream learning modules, limiting their ability to jointly improve trajectory generation and planning feasibility. To address these issues, we propose a hybrid trajectory planning framework that integrates IL-based multimodal trajectory proposal with differentiable optimization. In the proposed framework, an IL backbone generates candidate ego trajectories and surrounding-agent predictions, while a differentiable optimizer refines the selected trajectory using multi-objective cost functions with learnable weights related to safety, efficiency, and comfort. This design enables optimization objectives and constraints to provide gradient feedback to the upstream planning network, improving the consistency between candidate generation and downstream planning objectives. In addition, we introduce a surrounding agent centric data augmentation strategy that reuses real-world trajectories of surrounding vehicles as additional expert demonstrations, thereby enriching complex interaction and long-tail scenarios without extra data collection. Closed-loop experiments on the nuPlan benchmark show that the proposed method achieves a composite score of 94.04, outperforming PLUTO’s 93.14 while using only 30% of the training data. The results demonstrate that the proposed framework improves closed-loop planning performance, trajectory feasibility, and data efficiency under complex urban driving scenarios.
Get full access to this article
View all access options for this article.
