CrowdLLM: Building LLM-Based Digital Populations Augmented with Generative Models

Published Online:https://doi.org/10.1287/ijds.2025.0140

References

  • Aamari E, Kim J, Chazal F, Michel B, Rinaldo A, Wasserman L (2019) Estimating the reach of a manifold. Electronic J. Statist. 13(1):1359–1399.Google Scholar
  • Achiam J, Adler S, Agarwal S, Ahmad L, Akkaya I, Aleman FL, Almeida D, et al. (2023) GPT-4 technical report. Preprint, submitted March 15, https://arxiv.org/abs/2303.08774v1.Google Scholar
  • Alizadeh M, Gilardi F, Samei Z, Mosleh M (2025) Web-browsing LLMs can access social media profiles and infer user demographics. Preprint, submitted July 16, https://arxiv.org/abs/2507.12372.Google Scholar
  • Angelopoulos AN, Bates S, Fannjiang C, Jordan MI, Zrnic T (2023) Prediction-powered inference. Science 382(6671):669–674.Google Scholar
  • Anthis JR, Liu R, Richardson SM, Kozlowski AC, Koch B, Brynjolfsson E, Evans J, Bernstein MS (2025) Position: LLM social simulations are a promising research method. Proc. 42nd Internat. Conf. Machine Learn. (PMLR, New York), 81005–81034.Google Scholar
  • Ball S, Allmendinger S, Kreuter F, Kühl N (2025) Human preferences in large language model latent space: A technical analysis on the reliability of synthetic data in voting outcome prediction. Preprint, submitted February 22, https://arxiv.org/abs/2502.16280.Google Scholar
  • Binz M, Akata E, Bethge M, Brändle F, Callaway F, Coda-Forno J, Dayan P, et al. (2024) Centaur: A foundation model of human cognition. Preprint, submitted October 26, https://arxiv.org/abs/2410.20268v1.Google Scholar
  • Blin JM, Whinston AB (1975) Discriminant functions and majority voting. Management Sci. 21(5):557–566.LinkGoogle Scholar
  • Bui N, Nguyen HT, Kumar S, Theodore J, Qiu W, Nguyen VA, Ying R (2025) Mixture-of-personas language models for population simulation. Findings Assoc. Comput. Linguistics ACL 2025 (Association for Computational Linguistics, Stroudsburg, PA), 24761–24778.Google Scholar
  • Cai L, He J, Li Y, Liang J, Lin Y, Quan Z, Zeng Y, Xu J (2025) RTBAgent: A LLM-based agent system for real-time bidding. Comp. Proc. ACM Web Conf. 2025 (Association for Computing Machinery, New York), 104–113.Google Scholar
  • Cen SH, Ilyas A, Driss H, Park C, Hopkins A, Podimata C, et al. (2025) Large-scale, longitudinal study of large language models during the 2024 US election season. Preprint, submitted September 22, https://arxiv.org/abs/2509.18446.Google Scholar
  • Chang TY, Jia R (2022) Data curation alone can stabilize in-context learning. Preprint, submitted December 20, https://arxiv.org/abs/2212.10378v1.Google Scholar
  • Chen C, Yao B, Ye Y, Wang D, Li TJJ (2024) Evaluating the LLM agents for simulating humanoid behavior. Proc. 2024 CHI Workshop Human-Centered Evaluation Auditing Language Models (Association for Computing Machinery, New York).Google Scholar
  • Chen J, Gao C, Yuan S, Liu S, Cai Q, Jiang P (2025) Dlcrec: A novel approach for managing diversity in LLM-based recommender systems. Proc. 18th ACM Internat. Conf. Web Search Data Mining (Association for Computing Machinery, New York), 857–865.Google Scholar
  • Corecco N, Piatti G, Lanzendörfer LA, Fan FX, Wattenhofer R (2024) SUBER: An RL environment with simulated human behavior for recommender systems. Preprint, submitted August 20, https://arxiv.org/abs/2406.01631.Google Scholar
  • Costabile L, Orlando GM, La Gatta V, Moscato V (2025) Assessing the potential of generative agents in crowdsourced fact-checking. Preprint, submitted October 25, https://arxiv.org/abs/2504.19940.Google Scholar
  • Dahal B, Havrilla A, Chen M, Zhao T, Liao W (2022) On deep generative models for approximation and estimation of distributions on manifolds. Adv. Neural Inform. Processing Systems 35:10615–10628.Google Scholar
  • Deshmukh N, Venkatesh AS, Mathew A, Madugula M, Merchant PM, Lanham MA, Shirodkar S (2024) Harnessing LLMs to build an autonomous marketing agent. World Congress Comput. Sci. Comput. Engrg. Appl. Comput. (Springer, Cham, Switzerland), 271–282.Google Scholar
  • Dieng AB, Kim Y, Rush AM, Blei DM (2019) Avoiding latent variable collapse with generative skip models. 22nd Internat. Conf. Artificial Intelligence Statist. (PMLR, New York), 2397–2405.Google Scholar
  • Ding N, Qin Y, Yang G, Wei F, Yang Z, Su Y, Hu S, et al. (2023) Parameter-efficient fine-tuning of large-scale pre-trained language models. Nat. Machine Intelligence 5(3):220–235.Google Scholar
  • Dong YR, Hu T, Collier N (2024) Can LLM be a personalized judge? Preprint, submitted June 17, https://arxiv.org/abs/2406.11657.Google Scholar
  • Dubey A, Jauhri A, Pandey A, Kadian A, Al-Dahle A, Letman A, Mathur A, et al. (2024) The Llama 3 herd of models. Preprint, submitted November 23, https://arxiv.org/abs/2407.21783.Google Scholar
  • Eigner E, Händler T (2024) Determinants of LLM-assisted decision-making. Preprint, submitted February 27, https://arxiv.org/abs/2402.17385.Google Scholar
  • Gao Y, Lee D, Burtch G, Fazelpour S (2025) Take caution in using LLMs as human surrogates. Proc. Natl. Acad. Sci. USA 122(24):e2501660122.Google Scholar
  • Gemini Team, Anil R, Borgeaud S, Alayrac JB, Yu J, Soricut R, Schalkwyk J, et al. (2023) Gemini: A family of highly capable multimodal models. Preprint, submitted December 19, https://arxiv.org/abs/2312.11805v1.Google Scholar
  • Gemma Team, Kamath A, Ferret J, Pathak S, Vieillard N, Merhej R, Perrin S, et al. (2025) Gemma 3 technical report. Preprint, submitted March 25, https://arxiv.org/abs/2503.19786.Google Scholar
  • Goodfellow IJ, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y (2014) Generative adversarial nets. Adv. Neural Inform. Processing Systems 27:2672–2680.Google Scholar
  • Grunde-McLaughlin M, Lam MS, Krishna R, Weld DS, Heer J (2025) Designing LLM chains by adapting techniques from crowdsourcing workflows. ACM Trans. Comput.-Human Interaction 32(3):1–57.Google Scholar
  • Guo D, Yang D, Zhang H, Song J, Wang P, Zhu Q, Xu R, et al. (2025) DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645(8081):633–638.Google Scholar
  • Ho J, Jain A, Abbeel P (2020) Denoising diffusion probabilistic models. Adv. Neural Inform. Processing Systems 33:6840–6851.Google Scholar
  • Hou Y, Li J, He Z, Yan A, Chen X, McAuley J (2024) Bridging language and items for retrieval and recommendation. Preprint, submitted March 6, https://arxiv.org/abs/2403.03952.Google Scholar
  • Howe J (2006) The rise of crowdsourcing. Wired Magazine 14(6):176–183.Google Scholar
  • Hu T, Collier N (2024) Quantifying the persona effect in LLM simulations. Proc. 62nd Annu. Meeting Assoc. Comput. Linguistics, vol. 1 (Association for Computational Linguistics, Kerrville, TX), 10289–10307.Google Scholar
  • Hu EJ, Shen Y, Wallis P, Allen-Zhu Z, Li Y, Wang S, Wang L, Chen W (2021) LoRA: Low-rank adaptation of large language models. Preprint, submitted October 16, https://arxiv.org/abs/2106.09685.Google Scholar
  • Hua E, Qi B, Zhang K, Yu Y, Ding N, Lv X, Tian K, Zhou B (2024) Intuitive fine-tuning: Towards unifying SFT and RLHF into a single process. Preprint, submitted May 20, https://arxiv.org/abs/2405.11870v1.Google Scholar
  • Iren D, Bilgen S (2014) Cost of quality in crowdsourcing. Human Comput. 1(2):283–314.Google Scholar
  • Isinkaye FO, Folajimi YO, Ojokoh BA (2015) Recommendation systems: Principles, methods and evaluation. Egyptian Informatics J. 16(3):261–273.Google Scholar
  • Karten S, Li W, Ding Z, Kleiner S, Bai Y, Jin C (2025) LLM economist: Large population models and mechanism design in multi-agent generative simulacra. NeurIPS 2025 Workshop Algorithmic Collective Action (San Diego).Google Scholar
  • Kingma DP (2014) Adam: A method for stochastic optimization. Internat. Conf. Learn. Representations (San Diego).Google Scholar
  • Kingma DP, Welling M (2013) Auto-encoding variational Bayes. Preprint, submitted December 20, https://arxiv.org/abs/1312.6114v1.Google Scholar
  • Kirk R, Mediratta I, Nalmpantis C, Luketina J, Hambro E, Grefenstette E, Raileanu R (2023) Understanding the effects of RLHF on LLM generalisation and diversity. Preprint, submitted October 10, https://arxiv.org/abs/2310.06452v1.Google Scholar
  • Kobyzev I, Prince SJ, Brubaker MA (2021) Normalizing flows: An introduction and review of current methods. IEEE Trans. Pattern Anal. Machine Intelligence 43(11):3964–3979.Google Scholar
  • Leng Y, Yuan Y (2023) Do LLM agents exhibit social behavior? Preprint, submitted December 23, https://arxiv.org/abs/2312.15198v1.Google Scholar
  • Li J (2024a) A comparative study on annotation quality of crowdsourcing and LLM via label aggregation. ICASSP 2024 IEEE Internat. Conf. Acoustics Speech Signal Processing (IEEE, Piscataway, NJ), 6525–6529.Google Scholar
  • Li J (2024b) Human-LLM hybrid text answer aggregation for crowd annotations. Preprint, submitted October 22, https://arxiv.org/abs/2410.17099.Google Scholar
  • Li A, Chen H, Namkoong H, Peng T (2025) LLM generated persona is a promise with a catch. Preprint, submitted March 18, https://arxiv.org/abs/2503.16527.Google Scholar
  • Li Z, Li S, Wang Z, Lei N, Luo Z, Gu DX (2023) DPM-OT: A new diffusion probabilistic model based on optimal transport. Proc. IEEE/CVF Internat. Conf. Comput. Vision (IEEE, Piscataway, NJ), 22624–22633.Google Scholar
  • Li J, Zeng S, Wai HT, Li C, Garcia A, Hong M (2024) Getting more juice out of the SFT data: Reward learning from human demonstration improves SFT for LLM alignment. Adv. Neural Inform. Processing Systems 37:124292–124318. Google Scholar
  • Liu Z (2025) Cultural bias in large language models: A comprehensive analysis and mitigation strategies. J. Transcultural Commun. 3(2):224–244.Google Scholar
  • Liu O, Fu D, Yogatama D, Neiswanger W (2024a) DeLLMa: Decision making under uncertainty with large language models. Internat. Conf. Learn. Representations (Singapore).Google Scholar
  • Liu Y, Tao S, Zhao X, Zhu M, Ma W, Zhu J, Su C, et al. (2024b) CoachLM: Automatic instruction revisions improve the data quality in LLM instruction tuning. 2024 IEEE 40th Internat. Conf. Data Engrg. (IEEE, Piscataway, NJ), 5184–5197.Google Scholar
  • Liu Z, Chen C, Wang J, Chen M, Wu B, Che X, Wang D, Wang Q (2024c) Make LLM a testing expert: Bringing human-like interaction to mobile GUI testing via functionality-aware decisions. ICSE’24 Proc. IEEE/ACM 46th Internat. Conf. Software Engrg. (Association for Computing Machinery, New York), 1–13.Google Scholar
  • Lu Y, Huang J, Han Y, Bei B, Xie Y, Wang D, Wang J, He Q (2025) Beyond believability: Accurate human behavior simulation with fine-tuned LLMs. Preprint, submitted March 26, https://arxiv.org/abs/2503.20749v1.Google Scholar
  • Mai L, Carson-Berndsen J (2024) Improving linguistic diversity of large language models with possibility exploration fine-tuning. Preprint, submitted December 4, https://arxiv.org/abs/2412.03343.Google Scholar
  • Majumdar S, Elkind E, Pournaras E (2024) Generative AI voting: Fair collective choice is resilient to LLM biases and inconsistencies. Preprint, submitted May 31, https://arxiv.org/abs/2406.11871v1.Google Scholar
  • McAuley J, Targett C, Shi Q, Van Den Hengel A (2015) Image-based recommendations on styles and substitutes. Proc. 38th Internat. ACM SIGIR Conf. Res. Development Inform. Retrieval (Association for Computing Machinery, New York), 43–52.Google Scholar
  • Meng J (2024) AI emerges as the frontier in behavioral science. Proc. Natl. Acad. Sci. USA 121(10):e2401336121.Google Scholar
  • Moskovskiy D, Pletenev S, Panchenko A (2024) LLMs to replace crowdsourcing for parallel data creation? The case of text detoxification. Findings Assoc. Comput. Linguistics EMNLP 2024 (Association for Computational Linguistics, Kerrville, TX), 14361–14373.Google Scholar
  • Nussbaum Z, Morris JX, Duderstadt B, Mulyar A (2024) Nomic embed: Training a reproducible long context text embedder. Preprint, submitted February 2, https://arxiv.org/abs/2402.01613v1.Google Scholar
  • Olea C, Tucker H, Phelan J, Pattison C, Zhang S, Lieb M, Schmidt D, White J (2024) Evaluating persona prompting for question answering tasks. Internat. Conf. Artificial Intelligence Soft Comput.Google Scholar
  • Padmakumar V, He H (2023) Does writing with language models reduce content diversity? Internat. Conf. Learn. Representations (Vienna, Austria).Google Scholar
  • Pei J, Jurgens D (2023) When do annotator demographics matter? Measuring the influence of annotator demographics with the POPQUORN dataset. Proc. 17th Linguistic Annotation Workshop (Association for Computational Linguistics, Kerrville, TX), 252–265.Google Scholar
  • Penrose LS (1946) The elementary statistics of majority voting. J. Roy. Statist. Soc. 109(1):53–57.Google Scholar
  • Peterson AJ (2024) AI and the problem of knowledge collapse. Preprint, submitted April 22, https://arxiv.org/abs/2404.03502.Google Scholar
  • Piao J, Yan Y, Zhang J, Li N, Yan J, Lan X, Lu Z, et al. (2025) AgentSociety: Large-scale simulation of LLM-driven generative agents advances understanding of human behaviors and society. Preprint, submitted February 12, https://arxiv.org/abs/2502.08691v1.Google Scholar
  • Portugal IDS, Alencar P, Cowan D (2024) An agentic AI-based multi-agent framework for recommender systems. 2024 IEEE Internat. Conf. Big Data (IEEE, Piscataway, NJ), 5375–5382.Google Scholar
  • Rafailov R, Sharma A, Mitchell E, Manning CD, Ermon S, Finn C (2024) Direct preference optimization: Your language model is secretly a reward model. Adv. Neural Inform. Processing Systems 36:53728–53741.Google Scholar
  • Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O (2017) Proximal policy optimization algorithms. Preprint, submitted August 28, https://arxiv.org/abs/1707.06347.Google Scholar
  • Seedat N, Huynh N, Van Breugel B, Van Der Schaar M (2024) Curated LLM: Synergy of LLMs and data curation for tabular augmentation in low-data regimes. Proc. 41st Internat. Conf. Machine Learn. (PMLR, New York), 44060–44092.Google Scholar
  • Shu Y, Zhang H, Gu H, Zhang P, Lu T, Li D, Gu N (2024) RAH! RecSys–Assistant–Human: A human-centered recommendation framework with LLM agents. IEEE Trans. Comput. Soc. Systems 11(5):6759–6770.Google Scholar
  • Shypula A, Li S, Zhang B, Padmakumar V, Yin K, Bastani O (2025) Evaluating the diversity and quality of LLM generated content. Preprint, submitted April 16, https://arxiv.org/abs/2504.12522v1.Google Scholar
  • Sims MH, Hodges Shaw M, Gilbertson S, Storch J, Halterman MW (2019) Legal and ethical issues surrounding the use of crowdsourcing among healthcare providers. Health Informatics J. 25(4):1618–1630.Google Scholar
  • Song CH, Wu J, Washington C, Sadler BM, Chao WL, Su Y (2023) LLM-planner: Few-shot grounded planning for embodied agents with large language models. Proc. IEEE/CVF Internat. Conf. Comput. Vision (IEEE, Piscataway, NJ), 2998–3009.Google Scholar
  • Sun G, Zhan X, Such J (2024) Building better AI agents: A provocation on the utilisation of persona in LLM-based conversational agents. Proc. 6th ACM Conf. Conversational User Interfaces, vol. 35 (Association for Computing Machinery, New York), 1–6.Google Scholar
  • Tamura T, Ito H, Oyama S, Morishima A (2024) Simulation-based exploration for aggregation algorithms in human+AI crowd: What factors should we consider for better results? AAAI HCOMP 12th AAAI Conf. Human Comput. Crowdsourcing 2024 (AAAI Press, Washington, DC).Google Scholar
  • Touvron H, Lavril T, Izacard G, Martinet X, Lachaux MA, Lacroix T, Rozière B, et al. (2023) LLaMA: Open and efficient foundation language models. Preprint, submitted February 27, https://arxiv.org/abs/2302.13971.Google Scholar
  • Vaughan JW (2018) Making better use of the crowd: How crowdsourcing can advance machine learning research. J. Machine Learn. Res. 18(193):1–46.Google Scholar
  • Veselovsky V, Horta Ribeiro M, Cozzolino PJ, Gordon A, Rothschild D, West R (2025) Prevalence and prevention of large language model use in crowd work. Commun. ACM 68(3):42–47.Google Scholar
  • von der Heyde L, Haensch AC, Wenz A, Ma B (2024) United in diversity? Contextual biases in LLM-based predictions of the 2024 European Parliament elections. Preprint, submitted August 29, https://arxiv.org/abs/2409.09045v1.Google Scholar
  • Wang Q, Pan S, Linzen T, Black E (2025a) Multilingual prompting for improving LLM generation diversity. Preprint, submitted September 27, https://arxiv.org/abs/2505.15229.Google Scholar
  • Wang L, Gao H, Bo X, Chen X, Wen JR (2025b) YuLan-OneSim: Towards the next generation of social simulator with large language models. Preprint, submitted August 26, https://arxiv.org/abs/2505.07581.Google Scholar
  • Wang Y, Tang M, Shen N, Cui S, Wang W (2025c) Privacy risks of LLM-empowered recommender systems: An inversion attack perspective. Proc. 19th ACM Conf. Recommender Systems (Association for Computing Machinery, New York), 812–821.Google Scholar
  • Wang L, Zhang J, Yang H, Chen ZY, Tang J, Zhang Z, Chen X, et al. (2025d) User behavior simulation with large language model-based agents. ACM Trans. Inform. Systems 43(2):1–37.Google Scholar
  • Wang X, Wei J, Schuurmans D, Le QV, Chi EH, Narang S, Chowdhery A, Zhou D (2023) Self-consistency improves chain of thought reasoning in language models. Internat. Conf. Learn. Representations (Kigali, Rwanda).Google Scholar
  • Wei J, Wang X, Schuurmans D, Bosma M, Ichter B, Xia F, Chi EH, Le QV, Zhou D (2022) Chain-of-thought prompting elicits reasoning in large language models. Adv. Neural Inform. Processing Systems 35:24824–24837.Google Scholar
  • Whitehill J, Wu T-f, Bergsma J, Movellan J, Ruvolo P (2009) Whose vote should count more: Optimal integration of labels from labelers of unknown expertise. Adv. Neural Inform. Processing Systems 22:2035–2043.Google Scholar
  • Wood D, Mu T, Webb AM, Reeve HW, Luján M, Brown G (2023) A unified theory of diversity in ensemble learning. J. Machine Learn. Res. 24(359):1–49.Google Scholar
  • Wu X-K, Chen M, Li W, Wang R, Lu L, Liu J, Hwang K, et al. (2025) LLM fine-tuning: Concepts, opportunities, and challenges. Big Data Cognitive Comput. 9(4):87.Google Scholar
  • Wu T, Zhu H, Albayrak M, Axon A, Bertsch A, Deng W, Ding Z, et al. (2023) LLMs as workers in human-computational algorithms? replicating crowdsourcing pipelines with LLMs. Preprint, submitted July 19, https://arxiv.org/abs/2307.10168v1.Google Scholar
  • Xia H, McKernan B (2020) Privacy in crowdsourcing: A review of the threats and challenges. Comput. Supported Cooperative Work 29:263–301.Google Scholar
  • Xia Y, Kim J, Chen Y, Ye H, Kundu S, Hao C, Talati N (2024) Understanding the performance and estimating the cost of LLM fine-tuning. Preprint, submitted August 8, https://arxiv.org/abs/2408.04693.Google Scholar
  • Xie H, Maddalena E, Qarout R, Checco A (2023) The dark side of recruitment in crowdsourcing: Ethics and transparency in micro-task marketplaces. Comput. Supported Cooperative Work 32(3):439–474.Google Scholar
  • Xie C, Chen C, Jia F, Ye Z, Lai S, Shu K, Gu J, et al. (2024) Can large language model agents simulate human trust behavior? Adv. Neural Inform. Processing Systems 37:15674–15729.Google Scholar
  • Xu J, Han L, Sadiq S, Demartini G (2024a) On the role of large language models in crowdsourcing misinformation assessment. Proc. Internat. AAAI Conf. Web Soc. Media 18:1674–1686.Google Scholar
  • Xu W, Jojic N, Rao S, Brockett C, Dolan B (2024b) Echoes in AI: Quantifying lack of plot diversity in LLM outputs. Preprint, submitted December 31, https://arxiv.org/abs/2501.00273v1.Google Scholar
  • Yang JC, Dailisan D, Korecki M, Hausladen CI, Helbing D (2024a) LLM voting: Human choices and AI collective decision-making. Proc. AAAI/ACM Conf. AI Ethics Society, vol. 7 (AAAI Press, Washington, DC), 1696–1708.Google Scholar
  • Yang JC, Hausladen CI, Peters D, Pournaras E, Hnggli Fricker R, Helbing D (2024b) Designing digital voting systems for citizens: Achieving fairness and legitimacy in participatory budgeting. Digital Government Res. Practice 5(3):1–30.Google Scholar
  • Yang A, Li A, Yang B, Zhang B, Hui B, Zheng B, Yu B, et al. (2025) Qwen3 technical report. Preprint, submitted May 14, https://arxiv.org/abs/2505.09388.Google Scholar
  • Yeh M-H, Tao L, Wang J, Du X, Li Y (2024) How reliable is human feedback for aligning large language models? Preprint, submitted October 2, https://arxiv.org/abs/2410.01957v1.Google Scholar
  • Yin M, Zhou M (2018) Semi-implicit variational inference. Proc. 35th Internat. Conf. Machine Learn. (PMLR, New York), 5660–5669.Google Scholar
  • Yin F, Ye X, Durrett G (2024) LoFIT: Localized fine-tuning on LLM representations. Adv. Neural Inform. Processing Systems 37:9474–9506.Google Scholar
  • Zamfirescu-Pereira JD, Wong RY, Hartmann B, Yang Q (2023) Why Johnny can’t prompt: How non-AI experts try (and fail) to design LLM prompts. CHI’23 Proc. 2023 CHI Conf. Human Factors Comput. Systems (Association for Computing Machinery, New York), 1–21.Google Scholar
  • Zeng X, La Barbera D, Roitero K, Zubiaga A, Mizzaro S (2024) Combining large language models and crowdsourcing for hybrid human-AI misinformation detection. Proc. 47th International ACM SIGIR Conf. Res. Development Inform. Retrieval (Association for Computing Machinery, New York), 2332–2336.Google Scholar
  • Zhang Y, Zhou K, Liu Z (2024a) Neural prompt search. IEEE Trans. Pattern Anal. Machine Intelligence 47(7):5268–5280.Google Scholar
  • Zhang Y, Chen X, Zhou D, Jordan MI (2014) Spectral methods meet EM: A provably optimal algorithm for crowdsourcing. J. Machine Learn. Res. 17(102):1–44.Google Scholar
  • Zhang A, Chen Y, Sheng L, Wang X, Chua TS (2024b) On generative agents in recommendation. Proc. 47th Internat. ACM SIGIR Conf. Res. Development Inform. Retrieval (Association for Computing Machinery, New York), 1807–1817.Google Scholar
  • Zhang X, Lin J, Mou X, Yang S, Liu X, Sun L, Lyu H, et al. (2025a) SocioVerse: A world model for social simulation powered by LLM agents and a pool of 10 million real-world users. Preprint, submitted July 15, https://arxiv.org/abs/2504.10157.Google Scholar
  • Zhang Y, Diddee H, Holm S, Liu H, Liu X, Samuel V, Wang B, Ippolito D (2025b) NoveltyBench: Evaluating creativity and diversity in language models. Preprint, submitted August 9, https://arxiv.org/abs/2504.05228.Google Scholar
  • Zhang Z, Liu S, Liu Z, Zhong R, Cai Q, Zhao X, Zhang C, Liu Q, Jiang P (2025c) LLM-powered user simulator for recommender system. Proc. AAAI Conf. Artificial Intelligence, vol. 39 (AAAI Press, Washington, DC), 13339–13347.Google Scholar
  • Zhao X, Wang K, Peng W (2024) An electoral approach to diversify LLM-based multi-agent collective decision-making. Preprint, submitted October 19, https://arxiv.org/abs/2410.15168.Google Scholar
  • Zhao G, Yoon BJ, Park G, Jha S, Yoo S, Qian X (2025) Pareto prompt optimization. Internat. Conf. Learn. Representations (Singapore).Google Scholar
  • Zhou X, Zhu H, Mathur L, Zhang R, Yu H, Qi Z, Morency LP, et al. (2023) SOTOPIA: Interactive evaluation for social intelligence in language agents. Internat. Conf. Learn. Representations (Vienna, Austria).Google Scholar
INFORMS site uses cookies to store information on your computer. Some are essential to make our site work; Others help us improve the user experience. By using this site, you consent to the placement of these cookies. Please read our Privacy Statement to learn more.