Dispatch or Hold? An Inverse Optimization and Reinforcement Learning Approach for Multiobjective On-Demand Delivery

Published Online:https://doi.org/10.1287/trsc.2025.0149

In large-scale on-demand food delivery systems, dynamic order dispatching needs to balance immediate dispatch and order holding amid competing objectives, including delivery efficiency, service timeliness, and courier utilization. Holding certain orders for future consolidation may improve delivery efficiency but impair timeliness. In such systems, real-time decisions must be made to determine both when to dispatch orders and how to dispatch orders by grouping them to share dispatch routes. We propose a hierarchical offline estimation and on-policy learning framework that integrates inverse optimization with deep reinforcement learning. First, the framework decides which orders to dispatch immediately and which to hold for future periods. Next, it consolidates the dispatched orders by solving a multiobjective weighted set partitioning problem. We estimate the tradeoff weights across multiple delivery objectives using real-world data via cutting plane–based inverse optimization that supports combinatorial decisions. The learned cost structure is then embedded in a Markov decision process, and a dispatch policy is trained with proximal policy optimization to maximize long-term performance. We validate our approach using real-world data from a major food delivery platform. Compared with the platform’s current practice and the heuristic benchmarks, the learned policy significantly improves operational performance and provides the most effective balance across all operational objectives. The evaluation results demonstrate that strategic order holding can effectively increase grouping opportunities and improve overall delivery efficiency. In periods and areas with low order arrival density, a longer holding time can yield better consolidation opportunities and improved delivery efficiency. However, orders with long travel distances are less suitable for holding due to limited potential for grouping.

History: This paper has been accepted for the Transportation Science Special Issue on The First INFORMS TSL Data-Driven Research Challenge.

Funding: This research was supported by the Deutsche Forschungsgemeinschaft as part of the following research group: Advanced Optimization in a Networked Economy [Grant GRK2201/277991500] and by the Cross-Disciplinary Research Fund (CDRF) from George Washington University.

Supplemental Material: The online appendix is available at https://doi.org/10.1287/trsc.2025.0149.

INFORMS site uses cookies to store information on your computer. Some are essential to make our site work; Others help us improve the user experience. By using this site, you consent to the placement of these cookies. Please read our Privacy Statement to learn more.