Dispatch or Hold? An Inverse Optimization and Reinforcement Learning Approach for Multiobjective On-Demand Delivery

Published Online:https://doi.org/10.1287/trsc.2025.0149

References

  • Aswani A, Shen ZJ, Siddiq A (2018) Inverse optimization with noisy data. Oper. Res. 66(3):870–892.LinkGoogle Scholar
  • Boute RN, Gijsbrechts J, Van Jaarsveld W, Vanvuchelen N (2022) Deep reinforcement learning for inventory control: A roadmap. Eur. J. Oper. Res. 298(2):401–412.CrossrefGoogle Scholar
  • Carlsson JG, Song S (2018) Coordinated logistics with a truck and a drone. Management Sci. 64(9):4052–4069.LinkGoogle Scholar
  • Chan TC, Mahmood R, Zhu IY (2025) Inverse optimization: Theory and applications. Oper. Res. 73(2):1046–1074.LinkGoogle Scholar
  • Chan TC, Craig T, Lee T, Sharpe MB (2014) Generalized inverse multiobjective optimization with application to cancer therapy. Oper. Res. 62(3):680–695.LinkGoogle Scholar
  • Chen M, Hu M (2024) Courier dispatch in on-demand delivery. Management Sci. 70(6):3789–3807.LinkGoogle Scholar
  • Chen X, Ulmer MW, Thomas BW (2022) Deep Q-learning for same-day delivery with vehicles and drones. Eur. J. Oper. Res. 298(3):939–952.CrossrefGoogle Scholar
  • Chen S, Yan Z, Lim YF (2024) Managing the personalized order-holding problem in online retailing. Manufacturing Service Oper. Management 26(1):47–65.LinkGoogle Scholar
  • Chen JF, Wang L, Sang H, Wu CG, Wang J (2025) A prior and posterior order postponement framework for the on-demand food delivery problem. IEEE Trans. Intelligent Transportation Systems 26(10):14879–14895.CrossrefGoogle Scholar
  • Chen X, Xiong G, Lv Y, Chen Y, Song B, Wang FY (2021) A collaborative communication-Qmix approach for large-scale networked traffic signal control. Chen Y, Zheng N, Sotelo MA, Barbat S, Li L, eds. Proc. 24th IEEE Internat. Conf. Intelligent Transportation Systems (IEEE, Piscataway, NJ), 3450–3455.Google Scholar
  • Crönert T, Martin L, Minner S, Tang CS (2024) Inverse optimization of integer programming games for parameter estimation arising from competitive retail location selection. Eur. J. Oper. Res. 312(3):938–953.CrossrefGoogle Scholar
  • Ehmke JF, Campbell AM, Urban TL (2015) Ensuring service levels in routing problems with time windows and stochastic travel times. Eur. J. Oper. Res. 240(2):539–550.CrossrefGoogle Scholar
  • Ernst D, Geurts P, Wehenkel L (2005) Tree-based batch mode reinforcement learning. J. Machine Learn. Res. 6(18):503–556.Google Scholar
  • Hendrycks D (2016) Gaussian error linear units (GELUs). Preprint, submitted June 27, https://arxiv.org/abs/1606.08415.Google Scholar
  • Huang Y, Zhao L, Powell WB, Tong Y, Ryzhov IO (2019) Optimal learning for urban delivery fleet allocation. Transportation Sci. 53(3):623–641.LinkGoogle Scholar
  • James J, Yu W, Gu J (2019) Online vehicle routing with neural combinatorial optimization and deep reinforcement learning. IEEE Trans. Intelligent Transportation Systems 20(10):3806–3817.CrossrefGoogle Scholar
  • Klapp MA, Erera AL, Toriello A (2018) The one-dimensional dynamic dispatch waves problem. Transportation Sci. 52(2):402–415.LinkGoogle Scholar
  • Lang MA, Cleophas C, Ehmke JF (2021) Multi-criteria decision making in dynamic slotting for attended home deliveries. Omega (Westport) 102(C):102305. Google Scholar
  • Levine S, Kumar A, Tucker G, Fu J (2020) Offline reinforcement learning: Tutorial, review, and perspectives on open problems. Preprint, submitted May 4, https://arxiv.org/abs/2005.01643.Google Scholar
  • Liang Y, Luo H, Duan H, Li D, Liao H, Feng J, Zhao J, et al. (2024) Meituan’s real-time intelligent dispatching algorithms build the world’s largest minute-level delivery network. INFORMS J. Appl. Anal. 54(1):84–101.LinkGoogle Scholar
  • Liu S, Luo Z (2023) On-demand delivery from stores: Dynamic dispatching and routing with random demand. Manufacturing Service Oper. Management 25(2):595–612.LinkGoogle Scholar
  • Liu S, He L, Max Shen ZJ (2021) On-time last-mile delivery: Order assignment with travel-time predictors. Management Sci. 67(7):4095–4119.LinkGoogle Scholar
  • Lyu G, Cheung WC, Teo CP, Wang H (2024) Multiobjective stochastic optimization: A case of real-time matching in ride-sourcing markets. Manufacturing Service Oper. Management 26(2):500–518.LinkGoogle Scholar
  • Özarık SS, da Costa P, Florio AM (2024) Machine learning for data-driven last-mile delivery optimization. Transportation Sci. 58(1):27–44.LinkGoogle Scholar
  • Qin Z, Tang X, Jiao Y, Zhang F, Xu Z, Zhu H, Ye J (2020) Ride-hailing order dispatching at Didi via reinforcement learning. INFORMS J. Appl. Anal. 50(5):272–286.LinkGoogle Scholar
  • Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O (2017) Proximal policy optimization algorithms. Preprint, submitted July 20, https://arxiv.org/abs/1707.06347.Google Scholar
  • Shahmoradi Z, Lee T (2022) Quantile inverse optimization: Improving stability in inverse linear programming. Oper. Res. 70(4):2538–2562.LinkGoogle Scholar
  • Shi C, Qi Z, Wang J, Zhou F (2024) Value enhancement of reinforcement learning via efficient and robust trust region optimization. J. Amer. Statist. Assoc. 119(547):2011–2025.CrossrefGoogle Scholar
  • Statista (2025) Online food delivery: Worldwide. Accessed March 9, 2025, https://www.statista.com/outlook/emo/online-food-delivery/worldwide.Google Scholar
  • Tian J, Jia H, Wang G, Huang Q, Wu R, Gao H, Liu C (2025) Optimal scheduling of shared autonomous electric vehicles with multi-agent reinforcement learning: A MAPPO-based approach. Neurocomputing (Amsterdam) 622(C):129343.CrossrefGoogle Scholar
  • TSL-Meituan (2024) TSL data-driven research challenge. Accessed November 1, 2024, https://connect.informs.org/tsl/tslresources/datachallenge.Google Scholar
  • Ulmer MW, Thomas BW, Campbell AM, Woyak N (2021) The restaurant meal delivery problem: Dynamic pickup and delivery with deadlines and random ready times. Transportation Sci. 55(1):75–100.LinkGoogle Scholar
  • Van Hasselt H, Guez A, Silver D (2016) Deep reinforcement learning with double Q-learning. Schuurmans D, Wellman MP, eds. Proc. 30th AAAI Conf. Artificial Intelligence (AAAI Press, Palo Alto, CA), 2094–2100.Google Scholar
  • Wagner L, Calvo E, Amorim P (2023) Better together! The consumer implications of delivery consolidation. Manufacturing Service Oper. Management 25(3):903–920.LinkGoogle Scholar
  • Wang L (2009) Cutting plane algorithms for the inverse mixed integer linear programming problem. Oper. Res. Lett. 37(2):114–116.CrossrefGoogle Scholar
  • Wang H (2019) Routing and scheduling for a last-mile transportation system. Transportation Sci. 53(1):131–147.LinkGoogle Scholar
  • Wang Y, Minner S (2024) Deep reinforcement learning for demand fulfillment in online retail. Internat. J. Production Econom. 269(C):109133.CrossrefGoogle Scholar
  • Wang H, Odoni A (2016) Approximating the performance of a “last mile” transportation system. Transportation Sci. 50(2):659–675.LinkGoogle Scholar
  • Wang Y, Wang T, Wang X, Deng Y, Cao L (2024) Data-driven order fulfillment consolidation for online grocery retailing. INFORMS J. Appl. Anal. 54(3):211–221.LinkGoogle Scholar
  • Wei K, Vaze V, Jacquillat A (2022) Transit planning optimization under ride-hailing competition and traffic congestion. Transportation Sci. 56(3):725–749.LinkGoogle Scholar
  • Yan Y, Chow AH, Ho CP, Kuo YH, Wu Q, Ying C (2022) Reinforcement learning for logistics and supply chain management: Methodologies, state of the art, and future opportunities. Transportation Res. Part E: Logist. Transportation Rev. 162(C):102712.CrossrefGoogle Scholar
  • Yu Y, Zhou Q, Yi S, Zheng H, Wang S, Hao J, He R, Sun Z (2021) Delay to group in food delivery system: A prediction approach. Huang DS, Jo KH, Li J, Gribova V, Hussain A, eds. Proc. 17th Internat. Conf. Intelligent Comput., Part II, Lecture Notes in Computer Science, vol. 12837 (Springer, Cham, Switzerland), 540–551.Google Scholar
  • Zattoni Scroccaro P, van Beek P, Esfahani PM, Atasoy B (2024) Inverse optimization for routing problems. Transportation Sci. 59(2):301–321.LinkGoogle Scholar
INFORMS site uses cookies to store information on your computer. Some are essential to make our site work; Others help us improve the user experience. By using this site, you consent to the placement of these cookies. Please read our Privacy Statement to learn more.