Comparing Exploration–Exploitation Strategies of LLMs and Humans: Insights from Standard Multi-Armed Bandit Experiments
Published Online:1 Sep 2026https://doi.org/10.1287/ijds.2025.0113
References
- (2021) Attention-deficit/hyperactivity disorder and the explore/exploit trade-off. Neuropsychopharmacology 46(3):614–621.Google Scholar
- (2023) Using large language models to simulate multiple humans and replicate human subject studies. Internat. Conf. Machine Learn. (PMLR), 337–371.Google Scholar
- (2017) Revealing neurocomputational mechanisms of reinforcement learning and decision-making with the hBayesDM package. Comput. Psychiatry 1:24–57.Google Scholar
- (2011) A model-based fMRI analysis with hierarchical Bayesian parameter estimation. J. Neurosci. Psych. Econom. 4(2):95–110.Google Scholar
- (2025) What is your AI agent buying? Evaluation, implications and emerging questions for agentic e-commerce. Preprint, submitted August 4, https://arxiv.org/abs/2508.02630.Google Scholar
- (2005) Optimal Filtering (Courier Corporation, Chelmsford, MA).Google Scholar
- (2023) Out of one, many: Using language models to simulate human samples. Political Anal. (Oxford) 31(3):337–351.Google Scholar
- (2025) AI–human hybrids for marketing research: Leveraging large language models (LLMs) as collaborators. J. Marketing 89(2):43–70.Google Scholar
- (2002) Finite-time analysis of the multiarmed bandit problem. Machine Learn. 47:235–256.Google Scholar
- (2015) Transcranial stimulation over frontopolar cortex elucidates the choice attributes and neural mechanisms used to resolve exploration–exploitation trade-offs. J. Neuroscience 35(43):14544–14556.Google Scholar
- (2023) Using cognitive psychology to understand GPT-3. Proc. Natl. Acad. Sci. USA 120(6):e2218523120.Google Scholar
- (2017) Reminders of past choices bias decisions for reward in humans. Nature Comm. 8(1):15958.Google Scholar
- (2023) Using GPT for market research. Harvard Business School Marketing Unit Working Paper (23-062), Harvard Business School, Boston, MA.Google Scholar
- (2012) Regret analysis of stochastic and nonstochastic multi-armed bandit problems. Foundations Trends Machine Learn. 5(1):1–122.Google Scholar
- (2021) Increased random exploration in schizophrenia is associated with inflammation. NPJ Schizophrenia 7(1):6.Google Scholar
- (2020a) Dopaminergic modulation of the exploration/exploitation trade-off in human decision-making. Elife 9:e51260.Google Scholar
- (2020b) Dopaminergic modulation of the exploration/exploitation trade-off in human decision-making. Data set. https://doi.org/10.5281/zenodo.3872973.Google Scholar
- (2007) Should I stay or should I go? How the human brain manages the trade-off between exploitation and exploration. Philos. Trans. Roy. Soc. B: Biol. Sciences 362(1481):933–942.Google Scholar
- (2025) Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. Preprint, submitted July 7, https://arxiv.org/abs/2507.06261.Google Scholar
- (2006) Cortical substrates for exploratory decisions in humans. Nature 441(7095):876–879.Google Scholar
- (2023) Can AI language models replace human participants? Trends Cogn. Sci. 27(7):597–600.Google Scholar
- (2025) A behavioral model for exploration vs. exploitation: Theoretical framework and experimental evidence. Proc. 26th ACM Conf. Econom. Comput., 88–88.Google Scholar
- (2018) Deconstructing the human algorithms for exploration. Cognition 173:34–42.Google Scholar
- (2018) Dopaminergic genes are associated with both directed and random exploration. Neuropsychologia 120:97–104.Google Scholar
- (2024) Frontiers: Can large language models capture human preferences? Marketing Sci. 43(4):709–722.Link, Google Scholar
- (2023) The challenge of using LLMs to simulate human behavior: A causal inference perspective. Preprint, submitted December 24, https://arxiv.org/abs/2312.15524.Google Scholar
- (2025) DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning. Preprint, submitted January 22, https://arxiv.org/abs/2501.12948.Google Scholar
- (2023) Reasoning with language model is planning with world model. Proc. 2023 Conf. Empirical Methods in Natural Language Processing (Association for Computational Linguistics, Stroudsburg, PA). Google Scholar
- (2025) Should you use your large language model to explore or exploit? Preprint, submitted January 31, https://arxiv.org/abs/2502.00225.Google Scholar
- (2023) Large language models as simulated economic agents: What can we learn from Homo silicus? Technical report, National Bureau of Economic Research.Google Scholar
- (2026) WESE: Weak exploration to strong exploitation for LLM agents. Sci. China Inform. Sci. 69(3):132104.Google Scholar
- (2022) Inner monologue: Embodied reasoning through planning with language models. Preprint, submitted July 12, https://arxiv.org/abs/2207.05608.Google Scholar
- (2024) Social science meets LLMs: How reliable are large language models in social simulations? Preprint, submitted October 30, https://arxiv.org/abs/2410.23426.Google Scholar
- (2010) The role of the noradrenergic system in the exploration–exploitation trade-off: A psychopharmacological study. Front. Human Neurosci. 4:170.Google Scholar
- (2024) Decision-making behavior evaluation framework for LLMs under uncertain context. Adv. Neural Inform. Processing Systems 37:113360–113382. Google Scholar
- (1960) A new approach to linear filtering and prediction problems. J. Basic Engrg. 82(1):35–45.Google Scholar
- (2025) A survey of frontiers in LLM reasoning: Inference scaling, learning to reason, and agentic systems. Preprint, submitted April 12, https://arxiv.org/abs/2504.09037.Google Scholar
- (2022) Large language models are zero-shot reasoners. Adv. Neural Inform. Processing Systems 35:22199–22213.Google Scholar
- (2024) Can large language models explore in-context? Adv. Neural Inform. Processing Systems 37:120124–120158. Google Scholar
- (1985) Asymptotically efficient adaptive allocation rules. Adv. Appl. Math. 6(1):4–22.Google Scholar
- (2020) Bandit Algorithms (Cambridge University Press, Cambridge, UK).Google Scholar
- (2022) Pre-trained language models for interactive decision-making. Adv. Neural Inform. Processing Systems 35:31199–31212.Google Scholar
- (2024) Reason for future, act for now: A principled architecture for autonomous LLM agents. Forty-First Internat. Conf. Machine Learn.Google Scholar
- (2024) Evolve: Evaluating and optimizing LLMs for exploration. Preprint, submitted October 8, https://arxiv.org/abs/2410.06238.Google Scholar
- (2025) Do LLM agents have regret? A case study in online learning and games. Proc. Internat. Conf. Learn. Representations (ICLR) (ICLR, Appleton, WI).Google Scholar
- (2023) Generative agents: Interactive simulacra of human behavior. Proc. 36th Annual ACM Sympos. User Interface Software Technol., 1–22.Google Scholar
- (2024) Generative agent simulations of 1,000 people. Working paper.Google Scholar
- (2023) Generalization to new sequential decision making tasks with in-context learning. Preprint, submitted December 6, https://arxiv.org/abs/2312.03801.Google Scholar
- (2025) Towards scientific intelligence: A survey of LLM-based scientific agents. Preprint, submitted March 31, https://arxiv.org/abs/2503.24047.Google Scholar
- (2023) In-context impersonation reveals large language models’ strengths and biases. Adv. Neural Inform. Processing Systems 36:72044–72057.Google Scholar
- (2020) Finding structure in multi-armed bandits. Cognitive Psych. 119:101261.Google Scholar
- (2024) On the decision-making abilities in role-playing using large language models. Preprint, submitted February 29, https://arxiv.org/abs/2402.18807.Google Scholar
- (2022) Lower levels of directed exploration and reflective thinking are associated with greater anxiety and depression. Front. Psychiatry 12:782136.Google Scholar
- . (1998) Reinforcement Learning: An Introduction, vol. 1 (MIT Press, Cambridge, MA).Google Scholar
- (1933) On the likelihood that one unknown probability exceeds another in view of the evidence of two samples. Biometrika 25(3/4):285–294.Google Scholar
- (2025) Twin-2k-500: A data set for building digital twins of over 2,000 people based on their answers to over 500 questions. Marketing Sci. 44(6):1446–1455.Abstract, Google Scholar
- (2009) Discrete Choice Methods with Simulation (Cambridge University Press, Cambridge, UK).Google Scholar
- (2017) Practical Bayesian model evaluation using leave-one-out cross-validation and WAIC. Statist. Comput. 27:1413–1432.Google Scholar
- (2024) Pareto smoothed importance sampling. J. Machine Learn. Res. 25(72):1–58.Google Scholar
- (2005) Multi-armed bandit algorithms and empirical evaluation. Gama J, Camacho R, Brazdil PB, Jorge AM, Torgo L, eds. Machine Learning: ECML 2005, Lecture Notes in Computer Science, vol. 3720 (Springer, Berlin, Heidelberg), 437–448.Google Scholar
- (2026) Large language models for market research: A data-augmentation approach. Marketing Sci. 45(4):728–751.Google Scholar
- (2022) Chain-of-thought prompting elicits reasoning in large language models. Adv. Neural Inform. Processing Systems 35:24824–24837.Google Scholar
- (2021) Balancing exploration and exploitation with information and randomization. Curr. Opinion Behav. Sci. 38:49–56.Google Scholar
- (2014) Humans use directed and random exploration to solve the explore–exploit dilemma. J. Experiment. Psych. General 143(6):2074–2081.Google Scholar
- (2024) Can large language model agents simulate human trust behavior? The Thirty-Eighth Annual Conf. Neural Inform. Processing Systems.Google Scholar
- (2025) Toward large reasoning models: A survey of reinforced reasoning with large language models. Patterns 6(10):101370.Google Scholar
- (2023) Tree of thoughts: Deliberate problem solving with large language models. Adv. Neural Inform. Processing Systems 36:11809–11822.Google Scholar
- (2022) Automatic chain of thought prompting in large language models. Preprint, submitted October 7, https://arxiv.org/abs/2210.03493.Google Scholar
- (2025) Navigating the exploitation-exploration tradeoff: An empirical study of resource allocation in research labs. Preprint, submitted August 31, https://doi.org/10.2139/ssrn.5428035.Google Scholar

