Comparing Exploration–Exploitation Strategies of LLMs and Humans: Insights from Standard Multi-Armed Bandit Experiments

Published Online:https://doi.org/10.1287/ijds.2025.0113

Large language models (LLMs) are increasingly used to simulate or automate human behavior in complex sequential decision-making settings. A natural question is then whether LLMs exhibit similar decision-making behavior to humans, and can achieve comparable (or superior) performance. In this work, we focus on the exploration-exploitation (E&E) tradeoff, a fundamental aspect of dynamic decision-making under uncertainty. We employ canonical multiarmed bandit (MAB) experiments introduced in the cognitive science and psychiatry literature to conduct a comparative study of the E&E strategies of LLMs, humans, and MAB algorithms. We use interpretable choice models to capture the E&E strategies of the agents and investigate how enabling thinking traces, through both prompting strategies and thinking models, shapes LLM decision-making. We find that enabling thinking in LLMs shifts their behavior toward more human-like behavior, characterized by a mix of random and directed exploration. In a simple stationary setting, thinking-enabled LLMs exhibit similar levels of random and directed exploration compared with humans. However, in more complex, nonstationary environments, LLMs struggle to match human adaptability, particularly in effective directed exploration, despite achieving similar regret in certain scenarios. Our findings highlight both the promise and limits of LLMs as simulators of human behavior and tools for automated decision-making and point to potential areas for improvement.

History: This paper has been accepted by Yu Ding for the Virtual Special Issue on GenAI etc. for Business Analytics.

Funding: The research of N. Chen was supported by the Natural Sciences and Engineering Research Council of Canada (NSERC) Discovery [Grant RGPIN-2020-04038], and the research of V. Sarhangian was supported by the Natural Sciences and Engineering Research Council of Canada (NSERC) Discovery [Grant RGPIN-2025-06550].

Data Ethics & Reproducibility Note: The code capsule is available at https://github.com/zzy620/LLM-exploration-exploitation and in the Supplemental Material to this article (available at https://doi.org/10.1287/ijds.2025.0113).

INFORMS site uses cookies to store information on your computer. Some are essential to make our site work; Others help us improve the user experience. By using this site, you consent to the placement of these cookies. Please read our Privacy Statement to learn more.