Mean-Field Games of Speedy Information Access with Observation Costs

Published Online:https://doi.org/10.1287/moor.2024.0511

References

  • [1] Abel AB, Eberly JC, Panageas S (2013) Optimal inattention to the stock market with information costs and transactions costs. Econometrica 81(4):1455–1481.CrossrefGoogle Scholar
  • [2] Adlakha S, Lall S, Goldsmith A (2008) Information state for Markov decision processes with network delays. 2008 47th IEEE Conf. Decision Control (IEEE, Piscataway, NJ), 3840–3847.Google Scholar
  • [3] Adlakha S, Lall S, Goldsmith A (2011) Networked Markov decision processes with delays. IEEE Trans. Automatic Control 57(4):1013–1018.CrossrefGoogle Scholar
  • [4] Altman E, Nain P (1992) Closed-loop control with delayed information. ACM SIGMETRICS Performance Evaluation Rev. 20(1):193–204.CrossrefGoogle Scholar
  • [5] Anahtarcı B, Kariksiz C, Saldi N (2023) Q-learning in regularized mean-field games. Dynam. Games Appl. 13(1):89–117.Google Scholar
  • [6] Bander JL, White C (1999) Markov decision processes with noise-corrupted and delayed state observations. J. Oper. Res. Soc. 50(6):660–668.CrossrefGoogle Scholar
  • [7] Bellinger C, Coles R, Crowley M, Tamblyn I (2021) Active measure reinforcement learning for observation cost minimization. Antonie L, Zadeh PM, eds. Proc. 34th Canadian Conf. Artificial Intelligence (Canadian Artificial Intelligence Association, Vancouver, British Columbia).Google Scholar
  • [8] Bellinger C, Drozdyuk A, Crowley M, Tamblyn I (2022) Balancing information with observation costs in deep reinforcement learning. Kiringa I, Gambs S, Kalala KH, eds. Proc. 35th Canadian Conf. Artificial Intelligence (Canadian Artificial Intelligence Association, Toronto, Ontario).Google Scholar
  • [9] Bismut JM (1978) An introductory approach to duality in optimal stochastic control. SIAM Rev. 20(1):62–78.CrossrefGoogle Scholar
  • [10] Bogachev VI, Ruas MAS (2007) Measure Theory, vol. 1 (Springer, Berlin, Heidelberg).CrossrefGoogle Scholar
  • [11] Bruder B, Pham H (2009) Impulse control problem on finite horizon with execution delay. Stochastic Processes Appl. 119(5):1436–1469.CrossrefGoogle Scholar
  • [12] Caines PE, Huang M, Malhamé RP (2006) Large population stochastic dynamic games: Closed-loop McKean-Vlasov systems and the Nash certainty equivalence principle. Comm. Inform. Systems 6(3):221–252.CrossrefGoogle Scholar
  • [13] Cardaliaguet P, Lehalle CA (2018) Mean field game of controls and an application to trade crowding. Math. Financial Econom. 12:335–363.CrossrefGoogle Scholar
  • [14] Carmona R, Delarue F (2013) Mean field forward-backward stochastic differential equations. Electron. Comm. Probab. 18:1–15.CrossrefGoogle Scholar
  • [15] Carmona R, Delarue F (2018) Probabilistic Theory of Mean Field Games with Applications I and II, Probability Theory and Stochastic Modelling, vol. 83 (Springer, Cham, Switzerland).CrossrefGoogle Scholar
  • [16] Cartea Á, Sánchez-Betancourt L (2023) Optimal execution with stochastic delay. Finance Stochastics 27(1):1–47.CrossrefGoogle Scholar
  • [17] Chen B, Xu M, Li L, Zhao D (2021) Delay-aware model-based reinforcement learning for continuous control. Neurocomputing 450:119–128.CrossrefGoogle Scholar
  • [18] Cont R, Kotlicki A, Xu R (2021) Modelling COVID-19 contagion: Risk assessment and targeted mitigation policies. Roy. Soc. Open Sci. 8(3):201535.CrossrefGoogle Scholar
  • [19] Cooper C, Hahi N (1971) An optimal stochastic control problem with observation cost. IEEE Trans. Automatic Control 16(2):185–189.CrossrefGoogle Scholar
  • [20] Cui K, Koeppl H (2021) Approximately solving mean field games via entropy-regularized deep reinforcement learning. Banerjee A, Fukumizu K, eds. Proc. 24th Internat. Conf. Artificial Intelligence Statist., vol. 130 (PMLR, New York), 1909–1917.Google Scholar
  • [21] Dembo A, Zeitouni O (2009) Large Deviations Techniques and Applications, Stochastic Modelling and Applied Probability, vol. 38 (Springer, Berlin, Heidelberg).Google Scholar
  • [22] Dieudonné J (1960) Foundations of Modern Analysis (Academic Press, New York).Google Scholar
  • [23] Geist M, Scherrer B, Pietquin O (2019) A theory of regularized Markov decision processes. Chaudhuri K, Salakhutdinov R, eds. Proc. 36th Internat. Conf. Machine Learn., vol. 97 (PMLR, New York), 2160–2169.Google Scholar
  • [24] Georgii HO (2011) Gibbs Measures and Phase Transitions (Walter de Gruyter, Berlin).CrossrefGoogle Scholar
  • [25] Gomes DA, Voskanyan VK (2016) Extended deterministic mean-field games. SIAM J. Control Optim. 54(2):1030–1055.CrossrefGoogle Scholar
  • [26] Gu Z, Zhang J (2026) On information controls. Preprint, submitted February 7, https://arxiv.org/abs/2602.07318.Google Scholar
  • [27] Guo N, Kostina V (2021) Optimal causal rate-constrained sampling for a class of continuous Markov processes. IEEE Trans. Inform. Theory 67(12):7876–7890.CrossrefGoogle Scholar
  • [28] Guo X, Hu A, Zhang J (2024) MF-OMO: An optimization formulation of mean-field games. SIAM J. Control Optim. 62(1):243–270.CrossrefGoogle Scholar
  • [29] Guo X, Hu A, Xu R, Zhang J (2019) Learning mean-field games. Wallach HM, Larochelle H, Beygelzimer A, d’Alché-Buc F, Fox EB, eds. Proc. 33rd Internat. Conf. Neural Inform. Processing Systems (Curran Associates Inc., Red Hook, NY), 4966–4976.Google Scholar
  • [30] Guo X, Hu A, Xu R, Zhang J (2023a) A general framework for learning mean-field games. Math. Oper. Res. 48(2):656–686.LinkGoogle Scholar
  • [31] Guo X, Hu A, Santamaria M, Tajrobehkar M, Zhang J (2023b) MFGLib: A library for mean field games. Preprint, submitted April 17, https://arxiv.org/abs/2304.08630.Google Scholar
  • [32] Hajek B, Mitzel K, Yang S (2008) Paging and registration in cellular networks: Jointly optimal policies and an iterative algorithm. IEEE Trans. Inform. Theory 54(2):608–622.CrossrefGoogle Scholar
  • [33] Hernández-Lerma O (1989) Adaptive Markov Control Processes, Applied Mathematical Sciences, vol. 83 (Springer, New York).CrossrefGoogle Scholar
  • [34] Hernández-Lerma O, Lasserre JB (1996) Discrete-Time Markov Control Processes, Applications of Mathematics, vol. 30 (Springer, New York).CrossrefGoogle Scholar
  • [35] Huang Y, Zhu Q (2021) Self-triggered Markov decision processes. 2021 60th IEEE Conf. Decision Control (IEEE, Piscataway, NJ), 4507–4514.Google Scholar
  • [36] Katsikopoulos K, Engelbrecht S (2003) Markov decision processes with delays and asynchronous cost collection. IEEE Trans. Automatic Control 48(4):568–574.CrossrefGoogle Scholar
  • [37] Kerimkulov B, Leahy JM, Siska D, Szpruch L, Zhang Y (2025) A Fisher–Rao gradient flow for entropy-regularised Markov decision processes in Polish spaces. Foundations Comput. Math., ePub ahead of print August 11, https://doi.org/10.1007/s10208-025-09729-3. CrossrefGoogle Scholar
  • [38] Krueger D, Leike J, Evans O, Salvatier J (2020) Active reinforcement learning: Observing rewards at a cost. Preprint, submitted November 13, https://arxiv.org/abs/2011.06709.Google Scholar
  • [39] Lasry JM, Lions PL (2007) Mean field games. Japan J. Math. 2(1):229–260.CrossrefGoogle Scholar
  • [40] Laurière M, Tangpi L (2022) Convergence of large population games to mean field games with interaction through the controls. SIAM J. Math. Anal. 54(3):3535–3574.CrossrefGoogle Scholar
  • [41] Laurière M, Perrin S, Geist M, Pietquin O (2022) Learning mean field games: A survey. Preprint, submitted May 25, https://arxiv.org/abs/2205.12944.Google Scholar
  • [42] Nath S, Baranwal M, Khadilkar H (2021) Revisiting state augmentation methods for reinforcement learning with stochastic delays. CIKM’21: Proc. 30th ACM Internat. Conf. Inform. Knowledge Management (Association for Computing Machinery, New York), 1346–1355.Google Scholar
  • [43] Nayyar A, Başar T, Teneketzis D, Veeravalli VV (2013) Optimal strategies for communication and remote estimation with an energy harvesting sensor. IEEE Trans. Automatic Control 58(9):2246–2260.CrossrefGoogle Scholar
  • [44] Øksendal B, Sulem A (2008) Optimal stochastic impulse control with delayed reaction. Appl. Math. Optim. 58(2):243–255.CrossrefGoogle Scholar
  • [45] Pérolat J, Perrin S, Elie R, Laurière M, Piliouras G, Geist M, Tuyls K, Pietquin O (2022) Scaling mean field games by online mirror descent. AAMAS’22: Proc. 21st Internat. Conf. Autonomous Agents Multiagent Systems (International Foundation for Autonomous Agents and Multiagent Systems, Richland, SC), 1028–1037.Google Scholar
  • [46] Perrin S, Pérolat J, Laurière M, Geist M, Elie R, Pietquin O (2020) Fictitious play for mean field games: Continuous time analysis and applications. Larochelle H, Ranzato M, Hadsell R, Balcan MF, Lin H, eds. NIPS’20: Proc. 34th Internat. Conf. Neural Inform. Processing Systems (Curran Associates Inc., Red Hook, NY), 13199–13213.Google Scholar
  • [47] Reisinger C, Tam J (2025) Markov decision processes with observation costs: Framework and computation with a penalty scheme. Math. Oper. Res. 50(2):1305–1332.LinkGoogle Scholar
  • [48] Rong X, Yang L, Chu H, Fan M (2020) Effect of delay in diagnosis on transmission of COVID-19. Math. Biosci. Engrg. 17(3):2725–2740.CrossrefGoogle Scholar
  • [49] Saldi N, Başar T, Raginsky M (2018) Markov–Nash equilibria in mean-field games with discounted cost. SIAM J. Control Optim. 56(6):4256–4287.CrossrefGoogle Scholar
  • [50] Saldi N, Başar T, Raginsky M (2019) Approximate Nash equilibria in partially observed stochastic games with mean-field interactions. Math. Oper. Res. 44(3):1006–1033.LinkGoogle Scholar
  • [51] Saldi N, Başar T, Raginsky M (2023) Partially observed discrete-time risk-sensitive mean field games. Dynam. Games Appl. 13(3):926–960.CrossrefGoogle Scholar
  • [52] Saporito YF, Zhang J (2019) Stochastic control with delayed information and related nonlinear master equation. SIAM J. Control Optim. 57(1):693–717.CrossrefGoogle Scholar
  • [53] Schuitema E, Buşoniu L, Babuška R, Jonker P (2010) Control delay in reinforcement learning for real-time dynamic systems: A memoryless approach. 2010 IEEE/RSJ Internat. Conf. Intelligent Robots Systems (IEEE, Piscataway, NJ), 3226–3231.Google Scholar
  • [54] Tzoumas V, Carlone L, Pappas GJ, Jadbabaie A (2020) LQG control and sensing co-design. IEEE Trans. Automatic Control 66(4):1468–1483.CrossrefGoogle Scholar
  • [55] Winkelmann S, Schütte, C, Kleist Mv (2014) Markov control processes with rare state observation: Theory and application to treatment scheduling in HIV–1. Comm. Math. Sci. 12(5):859–877.CrossrefGoogle Scholar
  • [56] Wong R (2001) Asymptotic Approximations of Integrals, Classics in Applied Mathematics (Society for Industrial and Applied Mathematics, Philadelphia).CrossrefGoogle Scholar
  • [57] Wu W, Arapostathis A (2008) Optimal sensor querying: Markovian and LQG models with controlled observations. IEEE Trans. Automatic Control 53(6):1392–1405.CrossrefGoogle Scholar
  • [58] Yoshioka H, Tsujimura M (2020) Analysis and computation of an optimality equation arising in an impulse control problem with discrete and costly observations. J. Comput. Appl. Math. 366:112399.CrossrefGoogle Scholar
  • [59] Yoshioka H, Tsujimura M, Hamagami K, Yoshioka Y (2020a) A hybrid stochastic river environmental restoration modeling with discrete and costly observations. Optimal Control Appl. Methods 41(6):1964–1994.CrossrefGoogle Scholar
  • [60] Yoshioka H, Yaegashi Y, Tsujimura M, Yoshioka Y (2021) Cost-efficient monitoring of continuous-time stochastic processes based on discrete observations. Appl. Stochastic Models Bus. Indust. 37(1):113–138.CrossrefGoogle Scholar
  • [61] Yoshioka H, Yoshioka Y, Yaegashi Y, Tanaka T, Horinouchi M, Aranishi F (2020b) Analysis and computation of a discrete costly observation model for growth estimation and management of biological resources. Comput. Math. Appl. 79(4):1072–1093.CrossrefGoogle Scholar
INFORMS site uses cookies to store information on your computer. Some are essential to make our site work; Others help us improve the user experience. By using this site, you consent to the placement of these cookies. Please read our Privacy Statement to learn more.