Decision-Driven Regularization: A Blended Model for Learning and Optimization

Published Online:https://doi.org/10.1287/ijoc.2024.0930

References

  • Agrawal A, Amos B, Barratt S, Boyd S, Diamond S, Kolter JZ (2019) Differentiable convex optimization layers. Wallach H, Larochelle H, Beygelzimer A, d’Alche-Buc F, Fox E, Garnett R, eds. Advances in Neural Information Processing Systems, vol. 32 (Curran Associates, Red Hook, NY), 9562–9574.Google Scholar
  • Amos B, Kolter JZ (2017) Optnet: Differentiable optimization as a layer in neural networks. Precup D, Teh YW, eds. Proc. 34th Internat. Conf. Machine Learn., Proceedings of Machine Learning Research, vol. 70 (PMLR, New York), 136–145.Google Scholar
  • Amos B, Xu L, Kolter JZ (2017) Input convex neural networks. Precup D, Teh YW, eds. Proc. 34th Internat. Conf. Machine Learn., Proceedings of Machine Learning Research, vol. 70 (PMLR, New York), 146–155.Google Scholar
  • Athey S, Tibshirani J, Wager S (2019) Generalized random forests. Ann. Statist. 47(2):1148–1178.CrossrefGoogle Scholar
  • Ban GY, El Karoui N, Lim AE (2018) Machine learning and portfolio optimization. Management Sci. 64(3):1136–1154.LinkGoogle Scholar
  • Ban G, Gallien J, Mersereau AJ (2019) Dynamic procurement of new products with covariate information: The residual tree method. Manufacturing Service Oper. Management 21(4):798–815.LinkGoogle Scholar
  • Ben-Tal A, Den Hertog D, De Waegenaere A, Melenberg B, Rennen G (2013) Robust solutions of optimization problems affected by uncertain probabilities. Management Sci. 59(2):341–357.LinkGoogle Scholar
  • Bengio Y (1997) Using a financial training criterion rather than a prediction criterion. Internat. J. Neural Systems 8(04):433–443.CrossrefGoogle Scholar
  • Berden S, Mahmutoğulları Aİ, Tsouros D, Guns T (2026) Solver-free decision-focused learning for linear optimization problems. Proc. 39th Annual Conf. Neural Inform. Processing Systems (Curran Associates Inc., Red Hook, NY), 158127–158145.Google Scholar
  • Bertsimas D, Copenhaver M (2018) Characterization of the equivalence of robustification and regularization in linear and matrix regression. Eur. J. Oper. Res. 270(3):931–942.CrossrefGoogle Scholar
  • Bertsimas D, Kallus N (2020) From predictive to prescriptive analytics. Management Sci. 66(3):1025–1044.LinkGoogle Scholar
  • Bertsimas D, Van Parys B (2022) Bootstrap robust prescriptive analytics. Math. Programming 195(1):39–78.CrossrefGoogle Scholar
  • Blanchet J, Kang Y, Murthy K (2019) Robust Wasserstein profile inference and applications to machine learning. J. Appl. Probab. 56(3):830–857.CrossrefGoogle Scholar
  • Chan TC, Mahmood R, Zhu IY (2025) Inverse optimization: Theory and applications. Oper. Res. 73(2):1046–1074.LinkGoogle Scholar
  • Chung TH, Rostami V, Bastani H, Bastani O (2022) Decision-aware learning for optimizing health supply chains. Preprint, submitted November 15, https://arxiv.org/abs/2211.08507.Google Scholar
  • Delage E, Ye Y (2010) Distributionally robust optimization under moment uncertainty with application to data-driven problems. Oper. Res. 58(3):595–612.LinkGoogle Scholar
  • den Hertog D, Postek K (2016) Bridging the gap between predictive and prescriptive analytics-new optimization methodology needed. Preprint, submitted December 9 https://optimization-online.org/2016/12/5779/.Google Scholar
  • Deng Y, Sen S (2017) Learning enabled optimization: Towards a fusion of statistical learning and stochastic programming. Preprint, submitted March 14, https://optimization-online.org/2017/03/5904/.Google Scholar
  • Donti P, Amos B, Kolter J (2017) Task-based end-to-end model learning in stochastic optimization. Guyon I, von Luxburg U, Bengio S, Wallach H, Fergus R, Vishwanathan S, Garnett R, eds. Advances in Neural Information Processing Systems, vol. 30 (Curran Associates, Red Hook, NY), 5484–5494.Google Scholar
  • El Balghiti O, Elmachtoub A, Grigas P, Tewari A (2019) Generalization bounds in the predict-then-optimize framework. Wallach H, Larochelle H, Beygelzimer A, d’Alche-Buc F, Fox E, Garnett R, eds. Advances in Neural Information Processing Systems, vol. 32 (Curran Associates, Red Hook, NY), 14412–14421.Google Scholar
  • Elmachtoub A, Grigas P (2022) Smart “predict, then optimize.” Management Sci. 68(1):9–26.LinkGoogle Scholar
  • Elmachtoub AN, Lam H, Zhang H, Zhao Y (2023) Estimate-then-optimize versus integrated-estimation-optimization versus sample average approximation: A stochastic dominance perspective. Preprint, submitted April 13, https://arxiv.org/abs/2304.06833.Google Scholar
  • Esteban-Pérez A, Morales JM (2022) Distributionally robust stochastic programs with side information based on trimmings. Math. Programming 195(1):1069–1105.CrossrefGoogle Scholar
  • Estes AS, Richard JPP (2023) Smart predict-then-optimize for two-stage linear programs with side information. INFORMS J. Optim. 5(3):295–320.LinkGoogle Scholar
  • Ferreira K, Lee B, Simchi-Levi D (2016) Analytics for an online retailer: Demand forecasting and price optimization. Manufacturing Service Oper. Management 18(1):69–88.LinkGoogle Scholar
  • Gao R, Chen X, Kleywegt AJ (2017) Wasserstein distributional robustness and regularization in statistical learning. Preprint, submitted December 17, https://arxiv.org/abs/1712.06050.Google Scholar
  • Hannah L, Powell W, Blei D (2010) Nonparametric density estimation for stochastic optimization with an observable state variable. Lafferty J, Williams CKI, Shawe-Taylor J, Zemel R, Culotta A, eds. Advances in Neural Information Processing Systems, vol. 23 (Curran Associates, Red Hook, NY), 820–828.Google Scholar
  • Ho-Nguyen N, Kılınç-Karzan F (2022) Risk guarantees for end-to-end prediction and optimization processes. Management Sci. 68(12):8680–8698.LinkGoogle Scholar
  • Hu Y, Kallus N, Mao X (2022) Fast rates for contextual linear optimization. Management Sci. 68(6):4236–4245.LinkGoogle Scholar
  • Huang M, Gupta V (2024) Decision-focused learning with directional gradients. Globerson A, Mackey L, Belgrave D, Fan A, Paquet U, Tomczak J, Zhang C, eds. Advances in Neural Information Processing Systems, vol. 37 (NeurIPS) (Curran Associates, Red Hook, NY), 79194–79220.Google Scholar
  • Jeong J, Jaggi P, Butler A, Sanner S (2022) An exact symbolic reduction of linear smart predict+ optimize to mixed integer linear programming. Chaudhuri K, Jegelka S, Song L, Szepesvari C, Niu G, Sabato S, eds. Proc. 39th Internat. Conf. Machine Learn., Proceedings of Machine Learning Research, vol. 162 (PMLR, New York), 10053–10067.Google Scholar
  • Kallus N, Mao X (2023) Stochastic optimization forests. Management Sci. 69(4):1975–1994.LinkGoogle Scholar
  • Kannan R, Bayraksan G, Luedtke JR (2024) Residuals-based distributionally robust optimization with covariate information. Math. Programming 207(1):369–425.CrossrefGoogle Scholar
  • Kannan R, Bayraksan G, Luedtke JR (2025) Data-driven sample average approximation with covariate information. Oper. Res. 73(6):3245–3259.LinkGoogle Scholar
  • Kingma DP, Ba J (2015) Adam: A method for stochastic optimization. Proc. Internat. Conf. Learn. Representations (San Diego).Google Scholar
  • Kong L, Cui J, Zhuang Y, Feng R, Prakash BA, Zhang C (2022) End-to-end stochastic optimization with energy-based model. Koyejo S, Mohamed S, Agarwal A, Belgrave D, Cho K, Oh A, eds. Advances in Neural Information Processing Systems, vol. 35 (Curran Associates, Red Hook, NY):11341–11354.Google Scholar
  • Lawless C, Zhou A (2022) A note on task-aware loss via reweighing prediction loss by decision-regret. Preprint, submitted November 9, https://arxiv.org/abs/2211.05116.Google Scholar
  • Lin S, Chen Y, Li Y, Shen ZJM (2022) Data-driven newsvendor problems regularized by a profit risk constraint. Production Oper. Management 31(4):1630–1644.CrossrefGoogle Scholar
  • Liu S, He L, Shen Z (2021) On-time last mile delivery: Order assignment with travel time predictors. Management Sci. 67(7):4095–4119.LinkGoogle Scholar
  • Liyanage L, Shanthikumar G (2005) A practical inventory control policy using operational statistics. Oper. Res. Lett. 33(4):341–348.CrossrefGoogle Scholar
  • Loke GG, Tang Q, Xiao Y, Zhang X (2026) Decision-driven regularization: A blended model for learning and optimization. https://doi.org/10.1287/ijoc.2024.0930.cd, https://github.com/INFORMSJoC/2024.0930.Google Scholar
  • Mandi J, Stuckey P, Guns T (2020) Smart predict-and-optimize for hard combinatorial optimization problems. Proc. AAAI Conf. Artificial Intelligence, vol. 34 (AAAI Press, Palo Alto, CA), 1603–1610.Google Scholar
  • Mandi J, Bucarey V, Tchomba MMK, Guns T (2022) Decision-focused learning: Through the lens of learning to rank. Chaudhuri K, Jegelka S, Song L, Szepesvari C, Niu G, Sabato S, eds. Proc. 39th Internat. Conf. Machine Learn., Proceedings of Machine Learning Research, vol. 162 (PMLR, New York), 14935–14947.Google Scholar
  • Mandi J, Kotary J, Berden S, Mulamba M, Bucarey V, Guns T, Fioretto F (2024) Decision-focused learning: Foundations, state of the art, benchmark and future opportunities. J. Artificial Intelligence Res. 80:1623–1701.CrossrefGoogle Scholar
  • Mandi J, Mahmutogullari AI, Guns T (2025) Minimizing surrogate losses for decision-focused learning using differentiable optimization. Lynce I, Murano A, Vallati M, Villata S, Chesani F, Milano M, Omicini A, Dastani M, eds. 28th Eur. Conf. Artificial Intelligence, Including 14th Conf. Prestigious Appl. Intelligent Systems, Frontiers in Artificial Intelligence and Applications, vol. 413 (IOS Press, Amsterdam), 3888–3895.Google Scholar
  • McKenzie D, Wu_Fung S, Heaton H (2024) Differentiating through integer linear programs with quadratic regularization and Davis-Yin splitting. Trans. Machine Learn. Res.Google Scholar
  • Mohajerin Esfahani P, Kuhn D (2018) Data-driven distributionally robust optimization using the Wasserstein metric: Performance guarantees and tractable reformulations. Math. Programming 171(1):115–166.CrossrefGoogle Scholar
  • Mundru N (2019) Predictive and prescriptive methods in operations research and machine learning: An optimization approach. PhD thesis, Massachusetts Institute of Technology, Cambridge.Google Scholar
  • Muñoz MA, Pineda S, Morales JM (2022) A bilevel framework for decision-making under uncertainty with contextual information. Omega (Westport) 108:102575.CrossrefGoogle Scholar
  • Notz PM, Pibernik R (2022) Prescriptive analytics for flexible capacity management. Management Sci. 68(3):1756–1775.LinkGoogle Scholar
  • Perakis G, Roels G (2008) Regret in the newsvendor model with partial information. Oper. Res. 56(1):188–203.LinkGoogle Scholar
  • Perakis G, Sim M, Tang Q, Xiong P (2023) Robust pricing and production with information partitioning and adaptation. Management Sci. 69(3):1398–1419.LinkGoogle Scholar
  • Pogančić MV, Paulus A, Musil V, Martius G, Rolinek M (2019) Differentiation of black-box combinatorial solvers. Proc. Internat. Conf. Learn. Representations (ICLR).Google Scholar
  • Poursoltani M, Delage E (2022) Adjustable robust optimization reformulations of two-stage worst-case regret minimization problems. Oper. Res. 70(5):2906–2930.LinkGoogle Scholar
  • Qi M, Grigas P, Shen ZJ (2026) Integrated conditional estimation-optimization. Oper. Res. 74(3):1604–1625.LinkGoogle Scholar
  • Qi M, Shi Y, Qi Y, Ma C, Yuan R, Wu D, Shen ZM (2023) A practical end-to-end inventory management model with deep learning. Management Sci. 69(2):759–773.LinkGoogle Scholar
  • Sadana U, Chenreddy A, Delage E, Forel A, Frejinger E, Vidal T (2024) A survey of contextual optimization methods for decision-making under uncertainty. Eur. J. Oper. Res. 320(2):271–289.CrossrefGoogle Scholar
  • Siegel AF, Wagner MR (2021) Profit estimation error in the newsvendor model under a parametric demand distribution. Management Sci. 67(8):4863–4879.LinkGoogle Scholar
  • Sim M, Tang Q, Zhou M, Zhu T (2025) The analytics of robust satisficing: Predict, optimize, satisfice, then fortify. Oper. Res. 73(5):2708–2728.LinkGoogle Scholar
  • Sun C, Liu S, Li X (2023) Maximum optimality margin: A unified approach for contextual linear programming and inverse linear programming. Krause A, Brunskill E, Cho K, Engelhardt B, Sabato S, Scarlett J, eds. Proc. 40th Internat. Conf. Machine Learn., Proceedings of Machine Learning Research, vol. 202 (PMLR, New York), 32886–32912.Google Scholar
  • Tang B, Khalil EB (2024) Pyepo: A pytorch-based end-to-end predict-then-optimize library for linear and integer programming. Math. Programming Comput. 16(3):297–335.CrossrefGoogle Scholar
  • Tulabandhula T, Rudin C (2013) Machine learning with operational costs. J. Machine Learn. Res. 14(25):1989–2028.Google Scholar
  • Wang T, Chen N, Wang C (2024) Contextual optimization under covariate shift: A robust approach by intersecting Wasserstein balls. Preprint, submitted June 4, https://arxiv.org/abs/2406.02426.Google Scholar
  • Wang Y, Srivastava PR, Hanasusanto GA, Ho CP (2026) On data-driven prescriptive analytics with side information: A regularized Nadaraya–Watson approach. Manufacturing Service Oper. Management 28(3):841–859.LinkGoogle Scholar
  • Wilder B, Ewing E, Dilkina B, Tambe M (2019) End to end learning and optimization on graphs. Wallach H, Larochelle H, Beygelzimer A, d’Alche-Buc F, Fox E, Garnett R, eds. Advances in Neural Information Processing Systems, vol. 32 (Curran Associates, Red Hook, NY).Google Scholar
  • Wong E, Kolter Z (2018) Provable defenses against adversarial examples via the convex outer adversarial polytope. Dy J, Krause A, eds. Proc. 35th Internat. Conf. Machine Learn., Proceedings of Machine Learning Research, vol. 80 (PMLR, New York), 5286–5295.Google Scholar
  • Xu H, Caramanis C, Mannor S (2010) Robust regression and LASSO. IEEE Trans. Inform. Theory 56(7):3561–3574.CrossrefGoogle Scholar
  • Zhu T, Xie J, Sim M (2022) Joint estimation and robustness optimization. Management Sci. 68(3):1659–1677.LinkGoogle Scholar
INFORMS site uses cookies to store information on your computer. Some are essential to make our site work; Others help us improve the user experience. By using this site, you consent to the placement of these cookies. Please read our Privacy Statement to learn more.