Relaxed Indexability and Index Policy for Partially Observable Restless Bandits

Keqin Liu
Keqin Liu
[email protected]
https://orcid.org/0000-0002-4783-346X
School of Mathematics and Physics, Xi’an Jiaotong-Liverpool University, Suzhou 215123, China; and Jiangsu National Center for Applied Mathematics, Nanjing 210093, China
Search for more papers by this author

School of Mathematics and Physics, Xi’an Jiaotong-Liverpool University, Suzhou 215123, China; and Jiangsu National Center for Applied Mathematics, Nanjing 210093, China

Search for more papers by this author

Published Online:16 Apr 2025https://doi.org/10.1287/mnsc.2022.02831

Abstract

This paper addresses an important class of restless multiarmed bandit (RMAB) problems that finds broad application in operations research, stochastic optimization, and reinforcement learning. There are N independent Markov processes that may be operated, observed and offer rewards. Due to the resource constraint, we can only choose a subset of $M (M < N)$ processes to operate and accrue reward determined by the states of selected processes. We formulate the problem as a partially observable RMAB with an infinite state space and design an algorithm that achieves a near-optimal performance with low complexity. Our algorithm is based on a generalization of Whittle’s original idea of indexability. Referred to as the relaxed indexability, the extended definition leads to the efficient online verifications and computations of the approximate Whittle index under the proposed algorithmic framework.

This paper was accepted by Chung Piaw Teo, optimization.

Supplemental Material: The online appendix and data files are available at https://doi.org/10.1287/mnsc.2022.02831.

Volume 71, Issue 12

December 2025

Pages vii-x, 9869-10753, iv-vi

Article Information

Supplemental Material

Metrics

Information

Received:September 12, 2022
Accepted:December 20, 2024
Published Online:April 16, 2025

Cite as

Keqin Liu (2025) Relaxed Indexability and Index Policy for Partially Observable Restless Bandits. Management Science 71(12):10106-10121.

https://doi.org/10.1287/mnsc.2022.02831

Keywords

Acknowledgments

The author gratefully acknowledges the help from his students, Jiale Zha and Chengzhong Zhang, for the numerical analysis and figures. The author’s colleague Prof. Ting Wu and the anonymous reviewers provided very helpful comments for improving this paper.

PDF download

Available Issues

Available Issues

Available Issues

Available Issues

Available Issues

Available Issues

Available Issues

Available Issues

Available Issues

Relaxed Indexability and Index Policy for Partially Observable Restless Bandits

Abstract

Volume 71, Issue 12

Article Information

Supplemental Material

Metrics

Information

Cite as

Keywords

Sign Up for INFORMS Publications Updates and News