Top-Two Thompson Sampling for Contextual Selection Problems
Abstract
We aim to efficiently allocate a fixed simulation budget to identify the best designs for each context among a finite number of contexts. The performance of each design in a context is measured by an identifiable statistical characteristic, possibly with the existence of nuisance parameters. In a Bayesian framework, we extend the top-two Thompson sampling method designed for selecting the best design in a single context to the contextual selection problems, leading to an efficient sampling policy that simultaneously allocates simulation samples to both contexts and designs. To demonstrate the asymptotic optimality of the proposed sampling policy, we establish a posterior large deviations principle for the focal model, characterizing the exponential convergence rate of the posterior distribution for a broad range of identifiable sampling distribution families. The proposed sampling policy is then proved to be consistent and asymptotically satisfies a necessary and sufficient condition for optimality. We also discuss extending the proposed sampling policy to solve the contextual top- problems and show that a necessary condition for optimality is asymptotically attained. Numerical experiments demonstrate the good finite-sample performance of the proposed sampling policy.
Funding: This work was supported in part by the National Natural Science Foundation of China (NSFC) [Grants 72250065, 72022001, 71901003, 72325007, and 72501012], the China Postdoctoral Science Foundation [Grants 2025T180219 and 2025M770795], and the China Scholarship Council [Grant CSC202206010152]. This work was also supported in part by the Xiangjiang Laboratory Key Project [Grant 23XJ02004] and the Science and Technology Innovation Program of Hunan Province [Grant 2024RC7003].
Supplemental Material: The online appendix is available at https://doi.org/10.1287/moor.2023.0198.

