Comparing Exploration–Exploitation Strategies of LLMs and Humans: Insights from Standard Multi-Armed Bandit Experiments

Published Online:https://doi.org/10.1287/ijds.2025.0113

References

  • Addicott MA, Pearson JM, Schechter JC, Sapyta JJ, Weiss MD, Kollins SH (2021) Attention-deficit/hyperactivity disorder and the explore/exploit trade-off. Neuropsychopharmacology 46(3):614–621.Google Scholar
  • Aher GV, Arriaga RI, Kalai AT (2023) Using large language models to simulate multiple humans and replicate human subject studies. Internat. Conf. Machine Learn. (PMLR), 337–371.Google Scholar
  • Ahn WY, Haines N, Zhang L (2017) Revealing neurocomputational mechanisms of reinforcement learning and decision-making with the hBayesDM package. Comput. Psychiatry 1:24–57.Google Scholar
  • Ahn WY, Krawitz A, Kim W, Busemeyer JR, Brown JW (2011) A model-based fMRI analysis with hierarchical Bayesian parameter estimation. J. Neurosci. Psych. Econom. 4(2):95–110.Google Scholar
  • Allouah A, Besbes O, Figueroa JD, Kanoria Y, Kumar A (2025) What is your AI agent buying? Evaluation, implications and emerging questions for agentic e-commerce. Preprint, submitted August 4, https://arxiv.org/abs/2508.02630.Google Scholar
  • Anderson BD, Moore JB (2005) Optimal Filtering (Courier Corporation, Chelmsford, MA).Google Scholar
  • Argyle LP, Busby EC, Fulda N, Gubler JR, Rytting C, Wingate D (2023) Out of one, many: Using language models to simulate human samples. Political Anal. (Oxford) 31(3):337–351.Google Scholar
  • Arora N, Chakraborty I, Nishimura Y (2025) AI–human hybrids for marketing research: Leveraging large language models (LLMs) as collaborators. J. Marketing 89(2):43–70.Google Scholar
  • Auer P, Cesa-Bianchi N, Fischer P (2002) Finite-time analysis of the multiarmed bandit problem. Machine Learn. 47:235–256.Google Scholar
  • Beharelle AR, Polanía R, Hare TA, Ruff CC (2015) Transcranial stimulation over frontopolar cortex elucidates the choice attributes and neural mechanisms used to resolve exploration–exploitation trade-offs. J. Neuroscience 35(43):14544–14556.Google Scholar
  • Binz M, Schulz E (2023) Using cognitive psychology to understand GPT-3. Proc. Natl. Acad. Sci. USA 120(6):e2218523120.Google Scholar
  • Bornstein AM, Khaw MW, Shohamy D, Daw ND (2017) Reminders of past choices bias decisions for reward in humans. Nature Comm. 8(1):15958.Google Scholar
  • Brand J, Israeli A, Ngwe D (2023) Using GPT for market research. Harvard Business School Marketing Unit Working Paper (23-062), Harvard Business School, Boston, MA.Google Scholar
  • Bubeck S, Cesa-Bianchi N, et al. (2012) Regret analysis of stochastic and nonstochastic multi-armed bandit problems. Foundations Trends Machine Learn. 5(1):1–122.Google Scholar
  • Cathomas F, Klaus F, Guetter K, Chung HK, Raja Beharelle A, Spiller TR, Schlegel R, et al. (2021) Increased random exploration in schizophrenia is associated with inflammation. NPJ Schizophrenia 7(1):6.Google Scholar
  • Chakroun K, Mathar D, Wiehler A, Ganzer F, Peters J (2020a) Dopaminergic modulation of the exploration/exploitation trade-off in human decision-making. Elife 9:e51260.Google Scholar
  • Chakroun K, Mathar D, Wiehler A, Ganzer F, Peters J (2020b) Dopaminergic modulation of the exploration/exploitation trade-off in human decision-making. Data set. https://doi.org/10.5281/zenodo.3872973.Google Scholar
  • Cohen JD, McClure SM, Yu AJ (2007) Should I stay or should I go? How the human brain manages the trade-off between exploitation and exploration. Philos. Trans. Roy. Soc. B: Biol. Sciences 362(1481):933–942.Google Scholar
  • Comanici G, Bieber E, Schaekermann M, Pasupat I, Sachdeva N, Dhillon I, Blistein M, et al. (2025) Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. Preprint, submitted July 7, https://arxiv.org/abs/2507.06261.Google Scholar
  • Daw ND, O’doherty JP, Dayan P, Seymour B, Dolan RJ (2006) Cortical substrates for exploratory decisions in humans. Nature 441(7095):876–879.Google Scholar
  • Dillion D, Tandon N, Gu Y, Gray K (2023) Can AI language models replace human participants? Trends Cogn. Sci. 27(7):597–600.Google Scholar
  • Ding J, Feng Y, Rong Y (2025) A behavioral model for exploration vs. exploitation: Theoretical framework and experimental evidence. Proc. 26th ACM Conf. Econom. Comput., 88–88.Google Scholar
  • Gershman SJ (2018) Deconstructing the human algorithms for exploration. Cognition 173:34–42.Google Scholar
  • Gershman SJ, Tzovaras BG (2018) Dopaminergic genes are associated with both directed and random exploration. Neuropsychologia 120:97–104.Google Scholar
  • Goli A, Singh A (2024) Frontiers: Can large language models capture human preferences? Marketing Sci. 43(4):709–722.LinkGoogle Scholar
  • Gui G, Toubia O (2023) The challenge of using LLMs to simulate human behavior: A causal inference perspective. Preprint, submitted December 24, https://arxiv.org/abs/2312.15524.Google Scholar
  • Guo D, Yang D, Zhang H, Song J, Zhang R, Xu R, Zhu Q, et al. (2025) DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning. Preprint, submitted January 22, https://arxiv.org/abs/2501.12948.Google Scholar
  • Hao S, Gu Y, Ma H, Hong JJ, Wang Z, Wang DZ, Hu Z (2023) Reasoning with language model is planning with world model. Proc. 2023 Conf. Empirical Methods in Natural Language Processing (Association for Computational Linguistics, Stroudsburg, PA). Google Scholar
  • Harris K, Slivkins A (2025) Should you use your large language model to explore or exploit? Preprint, submitted January 31, https://arxiv.org/abs/2502.00225.Google Scholar
  • Horton JJ (2023) Large language models as simulated economic agents: What can we learn from Homo silicus? Technical report, National Bureau of Economic Research.Google Scholar
  • Huang X, Liu W, Chen X, Wang X, Lian D, Wang Y, Tang R, Chen E (2026) WESE: Weak exploration to strong exploitation for LLM agents. Sci. China Inform. Sci. 69(3):132104.Google Scholar
  • Huang W, Xia F, Xiao T, Chan H, Liang J, Florence P, Zeng A, et al. (2022) Inner monologue: Embodied reasoning through planning with language models. Preprint, submitted July 12, https://arxiv.org/abs/2207.05608.Google Scholar
  • Huang Y, Yuan Z, Zhou Y, Guo K, Wang X, Zhuang H, Sun W, et al. (2024) Social science meets LLMs: How reliable are large language models in social simulations? Preprint, submitted October 30, https://arxiv.org/abs/2410.23426.Google Scholar
  • Jepma M, Te Beek ET, Wagenmakers EJ, van Gerven JM, Nieuwenhuis S (2010) The role of the noradrenergic system in the exploration–exploitation trade-off: A psychopharmacological study. Front. Human Neurosci. 4:170.Google Scholar
  • Jia J, Yuan Z, Pan J, McNamara PE, Chen D (2024) Decision-making behavior evaluation framework for LLMs under uncertain context. Adv. Neural Inform. Processing Systems 37:113360–113382. Google Scholar
  • Kalman RE (1960) A new approach to linear filtering and prediction problems. J. Basic Engrg. 82(1):35–45.Google Scholar
  • Ke Z, Jiao F, Ming Y, Nguyen XP, Xu A, Long DX, Li M, et al. (2025) A survey of frontiers in LLM reasoning: Inference scaling, learning to reason, and agentic systems. Preprint, submitted April 12, https://arxiv.org/abs/2504.09037.Google Scholar
  • Kojima T, Gu SS, Reid M, Matsuo Y, Iwasawa Y (2022) Large language models are zero-shot reasoners. Adv. Neural Inform. Processing Systems 35:22199–22213.Google Scholar
  • Krishnamurthy A, Harris K, Foster DJ, Zhang C, Slivkins A (2024) Can large language models explore in-context? Adv. Neural Inform. Processing Systems 37:120124–120158. Google Scholar
  • Lai TL, Robbins H (1985) Asymptotically efficient adaptive allocation rules. Adv. Appl. Math. 6(1):4–22.Google Scholar
  • Lattimore T, Szepesvári C (2020) Bandit Algorithms (Cambridge University Press, Cambridge, UK).Google Scholar
  • Li S, Puig X, Paxton C, Du Y, Wang C, Fan L, Chen T, et al. (2022) Pre-trained language models for interactive decision-making. Adv. Neural Inform. Processing Systems 35:31199–31212.Google Scholar
  • Liu Z, Hu H, Zhang S, Guo H, Ke S, Liu B, Wang Z (2024) Reason for future, act for now: A principled architecture for autonomous LLM agents. Forty-First Internat. Conf. Machine Learn.Google Scholar
  • Nie A, Su Y, Chang B, Lee JN, Chi EH, Le QV, Chen M (2024) Evolve: Evaluating and optimizing LLMs for exploration. Preprint, submitted October 8, https://arxiv.org/abs/2410.06238.Google Scholar
  • Park C, Liu X, Ozdaglar A, Zhang K (2025) Do LLM agents have regret? A case study in online learning and games. Proc. Internat. Conf. Learn. Representations (ICLR) (ICLR, Appleton, WI).Google Scholar
  • Park JS, O’Brien J, Cai CJ, Morris MR, Liang P, Bernstein MS (2023) Generative agents: Interactive simulacra of human behavior. Proc. 36th Annual ACM Sympos. User Interface Software Technol., 1–22.Google Scholar
  • Park JS, Zou CQ, Shaw A, Hill BM, Cai C, Morris MR, Willer R, Liang P, Bernstein MS (2024) Generative agent simulations of 1,000 people. Working paper.Google Scholar
  • Raparthy SC, Hambro E, Kirk R, Henaff M, Raileanu R (2023) Generalization to new sequential decision making tasks with in-context learning. Preprint, submitted December 6, https://arxiv.org/abs/2312.03801.Google Scholar
  • Ren S, Jian P, Ren Z, Leng C, Xie C, Zhang J (2025) Towards scientific intelligence: A survey of LLM-based scientific agents. Preprint, submitted March 31, https://arxiv.org/abs/2503.24047.Google Scholar
  • Salewski L, Alaniz S, Rio-Torto I, Schulz E, Akata Z (2023) In-context impersonation reveals large language models’ strengths and biases. Adv. Neural Inform. Processing Systems 36:72044–72057.Google Scholar
  • Schulz E, Franklin NT, Gershman SJ (2020) Finding structure in multi-armed bandits. Cognitive Psych. 119:101261.Google Scholar
  • Shen C, Xie G, Zhang X, Xu J (2024) On the decision-making abilities in role-playing using large language models. Preprint, submitted February 29, https://arxiv.org/abs/2402.18807.Google Scholar
  • Smith R, Taylor S, Wilson RC, Chuning AE, Persich MR, Wang S, Killgore WD (2022) Lower levels of directed exploration and reflective thinking are associated with greater anxiety and depression. Front. Psychiatry 12:782136.Google Scholar
  • Sutton RS, Barto AG. (1998) Reinforcement Learning: An Introduction, vol. 1 (MIT Press, Cambridge, MA).Google Scholar
  • Thompson WR (1933) On the likelihood that one unknown probability exceeds another in view of the evidence of two samples. Biometrika 25(3/4):285–294.Google Scholar
  • Toubia O, Gui GZ, Peng T, Merlau DJ, Li A, Chen H (2025) Twin-2k-500: A data set for building digital twins of over 2,000 people based on their answers to over 500 questions. Marketing Sci. 44(6):1446–1455.AbstractGoogle Scholar
  • Train KE (2009) Discrete Choice Methods with Simulation (Cambridge University Press, Cambridge, UK).Google Scholar
  • Vehtari A, Gelman A, Gabry J (2017) Practical Bayesian model evaluation using leave-one-out cross-validation and WAIC. Statist. Comput. 27:1413–1432.Google Scholar
  • Vehtari A, Simpson D, Gelman A, Yao Y, Gabry J (2024) Pareto smoothed importance sampling. J. Machine Learn. Res. 25(72):1–58.Google Scholar
  • Vermorel J, Mohri M (2005) Multi-armed bandit algorithms and empirical evaluation. Gama J, Camacho R, Brazdil PB, Jorge AM, Torgo L, eds. Machine Learning: ECML 2005, Lecture Notes in Computer Science, vol. 3720 (Springer, Berlin, Heidelberg), 437–448.Google Scholar
  • Wang M, Zhang DJ, Zhang H (2026) Large language models for market research: A data-augmentation approach. Marketing Sci. 45(4):728–751.Google Scholar
  • Wei J, Wang X, Schuurmans D, Bosma M, Xia F, Chi E, Le QV, Zhou D (2022) Chain-of-thought prompting elicits reasoning in large language models. Adv. Neural Inform. Processing Systems 35:24824–24837.Google Scholar
  • Wilson RC, Bonawitz E, Costa VD, Ebitz RB (2021) Balancing exploration and exploitation with information and randomization. Curr. Opinion Behav. Sci. 38:49–56.Google Scholar
  • Wilson RC, Geana A, White JM, Ludvig EA, Cohen JD (2014) Humans use directed and random exploration to solve the explore–exploit dilemma. J. Experiment. Psych. General 143(6):2074–2081.Google Scholar
  • Xie C, Chen C, Jia F, Ye Z, Lai S, Shu K, Gu J, et al. (2024) Can large language model agents simulate human trust behavior? The Thirty-Eighth Annual Conf. Neural Inform. Processing Systems.Google Scholar
  • Xu F, Hao Q, Zong Z, Wang J, Zhang Y, Wang J, Lan X, et al. (2025) Toward large reasoning models: A survey of reinforced reasoning with large language models. Patterns 6(10):101370.Google Scholar
  • Yao S, Yu D, Zhao J, Shafran I, Griffiths T, Cao Y, Narasimhan K (2023) Tree of thoughts: Deliberate problem solving with large language models. Adv. Neural Inform. Processing Systems 36:11809–11822.Google Scholar
  • Zhang Z, Zhang A, Li M, Smola A (2022) Automatic chain of thought prompting in large language models. Preprint, submitted October 7, https://arxiv.org/abs/2210.03493.Google Scholar
  • Zhuo R (2025) Navigating the exploitation-exploration tradeoff: An empirical study of resource allocation in research labs. Preprint, submitted August 31, https://doi.org/10.2139/ssrn.5428035.Google Scholar
INFORMS site uses cookies to store information on your computer. Some are essential to make our site work; Others help us improve the user experience. By using this site, you consent to the placement of these cookies. Please read our Privacy Statement to learn more.