Nonstationary Experimental Design Under Structured Trends

Published Online:https://doi.org/10.1287/mnsc.2023.03329

References

  • Abadie A, Zhao J (2021) Synthetic controls for experimental design. Preprint, submitted August 4, https://arxiv.org/abs/2108.02196.Google Scholar
  • Abbasi-Yadkori Y, Pál D, Szepesvári C (2011 Improved algorithms for linear stochastic bandits. Shawe-Taylor J, Zemel R, Bartlett P, Pereira F, Weinberger K, eds. Advances in Neural Information Processing Systems, vol. 24 (Curran Associates Inc., Red Hook, NY), 2312–2320.Google Scholar
  • Adusumilli K (2025) Risk and optimal policies in bandit experiments. Econometrica 93(3):1003–1029.Google Scholar
  • Akaike H (2003) A new look at the statistical model identification. IEEE Trans. Automated Control 19(6):716–723.CrossrefGoogle Scholar
  • Akman VE, Raftery AE (1986) Asymptotic inference for a change-point Poisson process. Ann. Statist. 14(4):1583–1590.Google Scholar
  • Arora R, Marinov TV, Mohri M (2021) Corralling stochastic bandit algorithms. Proc. 24th Internat. Conf. Artificial Intelligence Statist., Proceedings of Machine Learning Research, vol. 130 (PMLR, New York), 2116–2124.Google Scholar
  • Atan O, Zame WR, Schaar M (2019) Sequential patient recruitment and allocation for adaptive clinical trials. Chaudhuri K, Sugiyama M, eds. Proc. 22nd Internat. Conf. Artificial Intelligence Statist., vol. 89 (PMLR, New York), 1891–1900.Google Scholar
  • Athey S, Eckles D, Imbens GW (2018) Exact p-values for network interference. J. Amer. Statist. Assoc. 113(521):230–240.CrossrefGoogle Scholar
  • Atsidakou A, Papadigenopoulos O, Basu S, Caramanis C, Shakkottai S (2021) Combinatorial blocking bandits with stochastic delays. Meila M, Zhang T, eds. Proc. 38th Internat. Conf. Machine Learn., vol. 139 (PMLR, New York), 404–413.Google Scholar
  • Auer P, Cesa-Bianchi N, Fischer P (2002a) Finite-time analysis of the multiarmed bandit problem. Machine Learn. 47(2–3):235–256.CrossrefGoogle Scholar
  • Auer P, Cesa-Bianchi N, Freund Y, Schapire RE (2002b) The nonstochastic multiarmed bandit problem. SIAM J. Comput. 32(1):48–77.CrossrefGoogle Scholar
  • Azevedo EM, Deng A, Montiel Olea JL, Rao J, Weyl EG (2020) A/b testing with fat tails. J. Political Econom. 128(12):4614–4000.CrossrefGoogle Scholar
  • Balsubramani A, Karnin Z, Schapire RE, Zoghi M (2016) Instance-dependent regret bounds for dueling bandits. Feldman V, Rakhlin A, Shamir O, eds. 29th Annual Conf. Learn. Theory, vol. 49 (PMLR, New York), 336–360.Google Scholar
  • Basu S, Papadigenopoulos O, Caramanis C, Shakkottai S (2021) Contextual blocking bandits. Banerjee A, Fukumizu K, eds. Proc. 24th Internat. Conf. Artificial Intelligence Statist., vol. 130 (PMLR, New York), 271–279.Google Scholar
  • Basu S, Sen R, Sanghavi S, Shakkottai S (2019) Blocking bandits. Wallach HM, Larochelle H, Beygelzimer A, d’Alché-Buc F, Fox EA, Garnett R, eds. Advances in Neural Information Processing Systems, vol. 32 (Curran Associates, Inc., Red Hook, NY), 4785–4794.Google Scholar
  • Berry DA (2006) Bayesian clinical trials. Nature Rev. Drug Discovery 5(1):27–36.CrossrefGoogle Scholar
  • Besbes O, Gur Y, Zeevi A (2014) Stochastic multi-armed-bandit problem with non-stationary rewards. Ghahramani Z, Welling M, Cortes C, Lawrence ND, Weinberger KQ, eds. Advances in Neural Information Processing Systems, vol. 27 (Curran Associates, Inc., Red Hook, NY), 199–207.Google Scholar
  • Besbes O, Gur Y, Zeevi A (2015) Non-stationary stochastic optimization. Oper. Res. 63(5):1227–1244.LinkGoogle Scholar
  • Besbes O, Gur Y, Zeevi A (2019) Optimal exploration–exploitation in a multi-armed bandit problem with non-stationary rewards. Stochastic Systems 9(4):319–337.LinkGoogle Scholar
  • Besson L, Kaufmann E, Maillard OA, Seznec J (2022) Efficient change-point detection for tackling piecewise-stationary bandits. J. Machine Learn. Res. 23(1):3337–3376.Google Scholar
  • Bhat N, Farias VF, Moallemi CC, Sinha D (2020) Near-optimal AB testing. Management Sci. 66(10):4477–4495.LinkGoogle Scholar
  • Bishop N, Chan H, Mandal D, Tran-Thanh L (2020) Adversarial blocking bandits. Larochelle H, Ranzato M, Hadsell R, Balcan MF, Lin H, eds. Adv. Neural Inform Processing Systems, vol. 33 (Curran Associates, Inc., Red Hook, NY), 8139–8149.Google Scholar
  • Bojinov I, Rambachan A, Shephard N (2021) Panel experiments and dynamic causal effects: A finite population perspective. Quant. Econom. 12(4):1171–1196.CrossrefGoogle Scholar
  • Bojinov I, Simchi-Levi D, Zhao J (2023) Design and analysis of switchback experiments. Management Sci. 69(7):3759–3777.LinkGoogle Scholar
  • Brennan J, Cong Y, Yu Y, Lin L, Peng Y, Meng C, Han N, Pouget-Abadie J, Holtz DM (2025) Reducing symbiosis bias through better A/B tests of recommendation algorithms. Proc. ACM Web Conf. (Association for Computing Machinery, New York), 3702–3715.Google Scholar
  • Cao Y, Wen Z, Kveton B, Xie Y (2019) Nearly optimal adaptive procedure with change detection for piecewise-stationary bandit. Chaudhuri K, Sugiyama M, eds. Proc. 22nd Internat. Conf. Artificial Intelligence Statist., vol. 89 (PMLR, New York), 418–427.Google Scholar
  • Chen N, Yang S, Zhang H (2025) Bridging adversarial and nonstationary multi-armed bandit. Production Oper. Management 34(8):2218–2231.Google Scholar
  • Chen N, Gao X, Xiong Y (2022) Debiasing samples from online learning using bootstrap. Camps-Valls G, Ruiz FJR, Valera I, eds. Proc. 25th Internat. Conf. Artificial Intelligence Statist. (PMLR, New York), 8514–8533.Google Scholar
  • Cheung WC, Simchi-Levi D, Zhu R (2022) Hedging the drift: Learning to optimize under nonstationarity. Management Sci. 68(3):1696–1713.LinkGoogle Scholar
  • Cheung WC, Simchi-Levi D, Zhu R (2023) Nonstationary reinforcement learning: The blessing of (more) optimism. Management Sci. 69(10):5722–5739.LinkGoogle Scholar
  • Chiang CK, Yang T, Lee CJ, Mahdavi M, Lu CJ, Jin R, Zhu S (2012) Online optimization with gradual variations. Mannor S, Srebro N, Williamson RC, eds. Proc. 25th Annual Conf. Learn. Theory, vol. 23 (PMLR, New York).Google Scholar
  • Chu W, Li L, Reyzin L, Schapire R (2011) Contextual bandits with linear payoff functions. Gordon G, Dunson D, Dudík M, eds. Proc. 14th Internat. Conf. Artificial Intelligence Statist., vol. 15 (PMLR, New York), 208–214.Google Scholar
  • Clerici G, Laforgue P, Cesa-Bianchi N (2023) Linear bandits with memory: From rotting to rising. Preprint, submitted February 16, https://arxiv.org/abs/2302.08345.Google Scholar
  • Dani V, Hayes TP, Kakade SM (2008) Stochastic linear optimization under bandit feedback. Servedio RA, Zhang T, eds. 21st Annual Conf. Learn. Theory (Omnipress, Madison, WI), 355–366.Google Scholar
  • Dimakopoulou M, Ren Z, Zhou Z (2021) Online multi-armed bandits with adaptive inference. Ranzato M, Beygelzimer A, Dauphin Y, Liang PS, Wortman Vaughan J, eds. Advances in Neural Information Processing Systems, vol. 34 (Curran Associates, Inc., Red Hook, NY), 1939–1951.Google Scholar
  • Edelman Y, Lee J (2025) AI capabilities progress has sped up. Accessed December 25, 2025, https://epoch.ai/data-insights/ai-capabilities-progress-has-sped-up.Google Scholar
  • Epoch AI (2025) Epoch capabilities index (ECI). Accessed December 25, 2025, https://epoch.ai/benchmarks/eci.Google Scholar
  • Farias VF, Li AA, Peng T, Zheng AT (2022a) Markovian interference in experiments. Proc. 36th Internat. Conf. Neural Inform. Processing Systems (Curran Associates Inc., Red Hook, NY).Google Scholar
  • Farias V, Moallemi C, Peng T, Zheng A (2022b) Synthetically controlled bandits. Preprint, submitted February 14, https://arxiv.org/abs/2202.07079.Google Scholar
  • Foster DJ, Rakhlin A, Simchi-Levi D, Xu Y (2020) Instance-dependent complexity of contextual bandits and reinforcement learning: A disagreement-based perspective. Preprint, submitted October 7, https://arxiv.org/abs/2010.03104.Google Scholar
  • Foussoul A, Goyal V, Gupta V (2023) MNL-bandit in non-stationary environments. Preprint, submitted March 4, https://arxiv.org/abs/2303.02504.Google Scholar
  • Garivier A, Moulines E (2011) On upper-confidence bound policies for switching bandit problems. Kivinen J, Szepesvári C, Ukkonen E, Zeugmann T, eds. Proc. Internat. Conf. Algorithmic Learn. Theory (Springer Berlin Heidelberg, Berlin, Heidelberg), 174–188.Google Scholar
  • Glynn PW, Zheng Z (2020) Estimation and inference for non-stationary arrival models with a linear trend. Mustafee N, Bae K-HG, Lazarova-Molnar S, Rabe M, Szabo C, Haas PJ, Son Y-J, eds. Proc. Winter Simulation Conf. (IEEE, Piscataway, NJ), 3764–3773.Google Scholar
  • Glynn PW, Johari R, Rasouli M (2020) Adaptive experimental design with temporal interference: A maximum likelihood approach. Larochelle H, Ranzato M, Hadsell R, Balcan MF, Lin H, eds. Advances in Neural Information Processing Systems, vol. 33 (Curran Associates, Inc., Red Hook, NY), 15054–15064.Google Scholar
  • Hahn J, Hirano K, Karlan D (2011) Adaptive experimental design using the propensity score. J. Bus. Econom. Statist. 29(1):96–108.CrossrefGoogle Scholar
  • Jackson PL, Muckstadt JA, Li Y (2019) Multiperiod stock allocation via robust optimization. Management Sci. 65(2):794–818.LinkGoogle Scholar
  • Jadbabaie A, Rakhlin A, Shahrampour S, Sridharan K (2015) Online optimization: Competing with dynamic comparators. Lebanon G, Vishwanathan SVN, eds. Proc. Eighteenth Artificial Intelligence Statist. (PMLR, New York), 398–406.Google Scholar
  • Johari R, Pekelis L, Walsh DJ (2015) Always valid inference: Bringing sequential analysis to a/b testing. Preprint, submitted December 15, https://arxiv.org/abs/1512.04922.Google Scholar
  • Johari R, Li H, Liskovich I, Weintraub GY (2022) Experimental design in two-sided platforms: An analysis of bias. Management Sci. 68(10):7069–7089.LinkGoogle Scholar
  • Karnin ZS, Anava O (2016) Multi-armed bandits: Competing with optimal sequences. Lee D, Sugiyama M, Luxburg U, Guyon I, Garnett R, eds. Advances in Neural Information Processing Systems, vol. 29 (Curran Associates, Inc., Red Hook, NY), 199–207.Google Scholar
  • Kasy M, Sautmann A (2021) Adaptive treatment assignment in experiments for policy choice. Econometrica 89(1):113–132.CrossrefGoogle Scholar
  • Kato M, Ishihara T, Honda J, Narita Y (2020) Efficient adaptive experimental design for average treatment effect estimation. Preprint, submitted February 13, https://arxiv.org/abs/2002.05308.Google Scholar
  • Keskin NB, Zeevi A (2017) Chasing demand: Learning and earning in a changing environment. Math. Oper. Res. 42(2):277–307.LinkGoogle Scholar
  • Kleinberg R, Immorlica N (2018) Recharging bandits. 2018 IEEE 59th Annual Sympos. Foundations Comput. Sci. (IEEE Computer Society, Washington, DC), 309–319.Google Scholar
  • Kohavi R, Henne RM, Sommerfield D (2007) Practical guide to controlled experiments on the web: Listen to your customers not to the hippo. Proc. 13th ACM SIGKDD Internat. Conf. Knowledge Discovery Data Mining (Association for Computing Machinery, New York), 959–967.Google Scholar
  • Kohavi R, Tang D, Xu Y (2020) Trustworthy Online Controlled Experiments: A Practical Guide to a/b Testing (Cambridge University Press, Cambridge, UK).CrossrefGoogle Scholar
  • Kuang X, Wager S (2024) Weak signal asymptotics for sequentially randomized experiments. Management Sci. 70(10):7024–7041.Google Scholar
  • Kuhl ME, Damerdji H, Wilson JR (1997) Estimating and simulating poisson processes with trends or asymmetric cyclic effects. Andradóttir S, Healy KJ, Withers DH, Nelson BL, eds. Proc. 1997 Winter Simulation Conf. (IEEE, Piscataway, NJ), 287–295.Google Scholar
  • Lai TL, Robbins H (1985) Asymptotically efficient adaptive allocation rules. Adv. Appl. Math. 6(1):4–22.CrossrefGoogle Scholar
  • Lai J, Xu L, Fang X, Dai T (2024) Regulating adaptive medical artificial intelligence: Can less oversight lead to greater compliance? Preprint, submitted November 21, https://doi.org/10.2139/ssrn.5009572.Google Scholar
  • Lattimore T, Szepesvári C (2020) Bandit Algorithms (Cambridge University Press, Cambridge, UK).CrossrefGoogle Scholar
  • Levine N, Crammer K, Mannor S (2017) Rotting bandits. Guyon I, Von Luxburg U, Bengio S, Wallach H, Fergus R, Vishwanathan S, Garnett R, eds. Advances in Neural Information Processing Systems, vol. 30 (Curran Associates, Inc., Red Hook, NY), 3075–3084.Google Scholar
  • Liu S, Jiang J, Li X (2022) Non-stationary bandits with knapsacks. Koyejo S, Mohamed S, Agarwal A, Belgrave D, Cho K, Oh A, eds. Advances in Neural Information Processing Systems, vol. 35 (Curran Associates, Inc., Red Hook, NY), 16522–16532. Google Scholar
  • Liu F, Lee J, Shroff N (2018) A change-detection based framework for piecewise-stationary multi-armed bandit problem. McIlraith SA, Weinberger KQ, eds. Proc. Thirty-Second AAAI Conf. Artificial Intelligence and Thirtieth Innovative Appl. Artificial Intelligence Conf. and Eighth AAAI Sympos. Educational Adv. Artificial Intelligence, vol. 32 (AAAI Press, Palo Alto, CA).Google Scholar
  • Liu Y, Van Roy B, Xu K (2023a) Nonstationary bandit learning via predictive sampling. Internat. Conf. Artificial Intelligence Statist. (PMLR, New York), 6215–6244.Google Scholar
  • Liu Y, Van Roy B, Xu K (2023b) A definition of non-stationary bandits. Preprint, submitted February 23, https://arxiv.org/abs/2302.12202.Google Scholar
  • Luo H, Wei CY, Agarwal A, Langford J (2018) Efficient contextual bandits in non-stationary worlds. Bubeck S, Perchet V, Rigollet P, eds. Proc. 31st Conf. Learn. Theory, vol. 75 (PMLR, New York), 1739–1776.Google Scholar
  • Mao W, Zhang K, Zhu R, Simchi-Levi D, Başar T (2025) Model-free nonstationary reinforcement learning: Near-optimal regret and applications in multiagent reinforcement learning and inventory control. Management Sci. 71(2):1564–1580.Google Scholar
  • Massey WA, Parker GA, Whitt W (1996) Estimating the parameters of a nonhomogeneous poisson process with linear rate. Telecomm. Systems 5:361–388.CrossrefGoogle Scholar
  • Metelli AM, Trovo F, Pirola M, Restelli M (2022) Stochastic rising bandits. Chaudhuri K, Jegelka S, Song L, Szepesvari C, Niu G, Sabato S, eds. Proc. 39th Internat. Conf. Machine Learn., vol. 162 (PMLR, New York), 15421–15457.Google Scholar
  • Min S, Russo D (2023) An information-theoretic analysis of nonstationary bandit learning. Internat. Conf. Machine Learn. (PMLR, New York), 24831–24849.Google Scholar
  • Mintz Y, Aswani A, Kaminsky P, Flowers E, Fukuoka Y (2020) Nonstationary bandits with habituation and recovery dynamics. Oper. Res. 68(5):1493–1516.LinkGoogle Scholar
  • Nie X, Tian X, Taylor J, Zou J (2018) Why adaptively collected data have negative bias and how to correct for it. Proc. Internat. Conf. Artificial Intelligence Statist. (PMLR), 1261–1269.Google Scholar
  • Piantadosi S (2017) Clinical Trials: A Methodologic Perspective (John Wiley & Sons, Hoboken, NJ).Google Scholar
  • Qin C, Russo D (2022) Adaptivity and confounding in multi-armed bandit experiments. Preprint, submitted XX, https://arxiv.org/abs/2202.09036.Google Scholar
  • Rosenthal JT, Beecy A, Sabuncu MR (2025) Rethinking clinical trials for medical ai with dynamic deployments of adaptive systems. NPJ Digital Medicine 8(1):1–6.CrossrefGoogle Scholar
  • Schwarz G (1978) Estimating the dimension of a model. Ann. Statist. 6(2):461–464.Google Scholar
  • Seznec J, Menard P, Lazaric A, Valko M (2020) A single algorithm for both restless and rested rotting bandits. Chiappa S, Calandra R, eds. Proc. Twenty Third Internat. Conf. Artificial Intelligence Statist., vol. 108 (PMLR, New York), 3784–3794.Google Scholar
  • Seznec J, Locatelli A, Carpentier A, Lazaric A, Valko M (2019) Rotting bandits are no harder than stochastic ones. Chaudhuri K, Sugiyama M, eds. Proc. 22nd Internat. Conf. Artificial Intelligence Statist., vol. 89 (PMLR, New York), 2564–2572.Google Scholar
  • Simchi-Levi D, Wang C (2022) Multi-armed bandit experimental design: Online decision-making and adaptive inference. Management Sci. 71(6):4828–4846.Google Scholar
  • Simon R (1977) Adaptive treatment assignment methods and clinical trials. Biometrics 33(4):743–749.CrossrefGoogle Scholar
  • Su Y, Wang X, Le EY, Liu L, Li Y, Lu H, Lipshitz B, et al. (2024) Long-term value of exploration: Measurements, findings and algorithms. Caudillo-Mata LA, Lattanzi S, Muñoz Medina A, Akoglu L, Gionis A, Vassilvitskii S, eds. Proc. 17th ACM Internat. Conf. Web Search Data Mining (Association for Computing Machinery, New York), 636–644.Google Scholar
  • Thompson WR (1933) On the likelihood that one unknown probability exceeds another in view of the evidence of two samples. Biometrika 25(3/4):285–294.CrossrefGoogle Scholar
  • Wager S, Xu K (2021) Experimenting in equilibrium. Management Sci. 67(11):6694–6715.LinkGoogle Scholar
  • Wainwright MJ (2019) High-Dimensional Statistics: A Non-Asymptotic Viewpoint, vol. 48 (Cambridge University Press, Cambridge, UK).CrossrefGoogle Scholar
  • Wang Y, Chen B, Simchi-Levi D (2021) Multimodal dynamic pricing. Management Sci. 67(10):6136–6152.LinkGoogle Scholar
  • Wang S, Huang L, Lui J (2020) Restless-UCB, an efficient and low-complexity algorithm for online restless bandits. Larochelle H, Ranzato M, Hadsell R, Balcan MF, Lin H, eds. Advances in Neural Information Processing Systems, vol. 33 (Curran Associates, Inc., Red Hook, NY), 11878–11889.Google Scholar
  • Wang J, Zhao P, Zhou ZH (2023) Revisiting weighted strategy for non-stationary parametric bandits. Preprint, submitted March 5, https://arxiv.org/abs/2303.02691.Google Scholar
  • Wei CY, Hong YT, Lu CJ (2016) Tracking the best expert in non-stationary stochastic environments. Lee D, Sugiyama M, Luxburg U, Guyon I, Garnett R, eds. Advances in Neural Information Processing Systems, vol. 29 (Curran Associates, Inc., Red Hook, NY), 3972–3980.Google Scholar
  • Whittle P (1988) Restless bandits: Activity allocation in a changing world. J. Appl. Probability 25(A):287–298.CrossrefGoogle Scholar
  • Wu Y, Zheng Z (2022) Performance evaluation and stochastic optimization with gradually changing non-stationary data. Preprint, submitted January 20, https://doi.org/10.2139/ssrn.4013301.Google Scholar
  • Wu Y, Zheng Z, Zhang G, Zhang Z, Wang C (2025) Nonstationary A/B tests: Optimal variance reduction, bias correction, and valid inference. Management Sci. 71(6):4707–4727.Google Scholar
  • Xie L, Zou S, Xie Y, Veeravalli VV (2021) Sequential (quickest) change detection: Classical results and new directions. IEEE J. Selected Areas Inform. Theory 2(2):494–514.CrossrefGoogle Scholar
  • Xiong R, Athey S, Bayati M, Imbens G (2024) Optimal experimental design for staggered rollouts. Management Sci. 70(8):5317–5336.Google Scholar
  • Xu X, Dong F, Li Y, He S, Li X (2020) Contextual-bandit based personalized recommendation with time-varying user interests. Conitzer V, Sha F, eds. Proc. Thirty-Fourth AAAI Conf. Artificial Intelligence, vol. 34 (AAAI Press, Palo Alto, CA), 6518–6525.Google Scholar
  • Yancey KP, Settles B (2020) A sleeping, recovering bandit algorithm for optimizing recurring notifications. Gupta R, Liu Y, Tang J, Aditya Prakash B, eds. Proc. 26th ACM SIGKDD Internat. Conf. Knowledge Discovery Data Mining (Association for Computing Machinery, New York), 3008–3016.Google Scholar
  • Zhang X, Frazier PI (2021) Restless bandits with many arms: Beating the central limit theorem. Preprint, submitted July 25, https://arxiv.org/abs/2107.11911.Google Scholar
  • Zhang K, Janson L, Murphy S (2020) Inference for batched bandits. Larochelle H, Ranzato M, Hadsell R, Balcan MF, Lin H, eds. Advances in Neural Information Processing Systems, vol. 33 (Curran Associates, Inc., Red Hook, NY), 9818–9829. Google Scholar
  • Zhao J, Zhou Z (2025) Pigeonhole design: Balancing sequential experiments from an online matching perspective. Management Sci. 71(3):1889–1908.Google Scholar
  • Zheng Z, Glynn PW (2017) Fitting continuous piecewise linear poisson intensities via maximum likelihood and least squares. Chan WKV, D’Ambrogio A, Zacharewicz G, Mustafee N, Wainer G, Page E, eds. Proc. 2017 Winter Simulation Conf. (IEEE, Piscataway, NJ), 1740–1749.Google Scholar
  • Zhou X, Xiong Y, Chen N, Gao X (2021) Regime switching bandits. Ranzato M, Beygelzimer A, Dauphin Y, Liang PS, Wortman Vaughan J, eds. Advances in Neural Information Processing Systems, vol. 34 (Curran Associates, Inc., Red Hook, NY), 4542–4554.Google Scholar
  • Zhu F, Zheng Z (2020) When demands evolve larger and noisier: Learning and earning in a growing environment. Daumé H III, Singh A, eds. Proc. 37th Internat. Conf. Machine Learn. (PMLR, New York), 11629–11638.Google Scholar
INFORMS site uses cookies to store information on your computer. Some are essential to make our site work; Others help us improve the user experience. By using this site, you consent to the placement of these cookies. Please read our Privacy Statement to learn more.