CrowdLLM: Building LLM-Based Digital Populations Augmented with Generative Models
Published Online:5 Aug 2026https://doi.org/10.1287/ijds.2025.0140
References
- (2019) Estimating the reach of a manifold. Electronic J. Statist. 13(1):1359–1399.Google Scholar
- (2023) GPT-4 technical report. Preprint, submitted March 15, https://arxiv.org/abs/2303.08774v1.Google Scholar
- (2025) Web-browsing LLMs can access social media profiles and infer user demographics. Preprint, submitted July 16, https://arxiv.org/abs/2507.12372.Google Scholar
- (2023) Prediction-powered inference. Science 382(6671):669–674.Google Scholar
- (2025) Position: LLM social simulations are a promising research method. Proc. 42nd Internat. Conf. Machine Learn. (PMLR, New York), 81005–81034.Google Scholar
- (2025) Human preferences in large language model latent space: A technical analysis on the reliability of synthetic data in voting outcome prediction. Preprint, submitted February 22, https://arxiv.org/abs/2502.16280.Google Scholar
- (2024) Centaur: A foundation model of human cognition. Preprint, submitted October 26, https://arxiv.org/abs/2410.20268v1.Google Scholar
- (1975) Discriminant functions and majority voting. Management Sci. 21(5):557–566.Link, Google Scholar
- (2025) Mixture-of-personas language models for population simulation. Findings Assoc. Comput. Linguistics ACL 2025 (Association for Computational Linguistics, Stroudsburg, PA), 24761–24778.Google Scholar
- (2025) RTBAgent: A LLM-based agent system for real-time bidding. Comp. Proc. ACM Web Conf. 2025 (Association for Computing Machinery, New York), 104–113.Google Scholar
- (2025) Large-scale, longitudinal study of large language models during the 2024 US election season. Preprint, submitted September 22, https://arxiv.org/abs/2509.18446.Google Scholar
- (2022) Data curation alone can stabilize in-context learning. Preprint, submitted December 20, https://arxiv.org/abs/2212.10378v1.Google Scholar
- (2024) Evaluating the LLM agents for simulating humanoid behavior. Proc. 2024 CHI Workshop Human-Centered Evaluation Auditing Language Models (Association for Computing Machinery, New York).Google Scholar
- (2025) Dlcrec: A novel approach for managing diversity in LLM-based recommender systems. Proc. 18th ACM Internat. Conf. Web Search Data Mining (Association for Computing Machinery, New York), 857–865.Google Scholar
- (2024) SUBER: An RL environment with simulated human behavior for recommender systems. Preprint, submitted August 20, https://arxiv.org/abs/2406.01631.Google Scholar
- (2025) Assessing the potential of generative agents in crowdsourced fact-checking. Preprint, submitted October 25, https://arxiv.org/abs/2504.19940.Google Scholar
- (2022) On deep generative models for approximation and estimation of distributions on manifolds. Adv. Neural Inform. Processing Systems 35:10615–10628.Google Scholar
- (2024) Harnessing LLMs to build an autonomous marketing agent. World Congress Comput. Sci. Comput. Engrg. Appl. Comput. (Springer, Cham, Switzerland), 271–282.Google Scholar
- (2019) Avoiding latent variable collapse with generative skip models. 22nd Internat. Conf. Artificial Intelligence Statist. (PMLR, New York), 2397–2405.Google Scholar
- (2023) Parameter-efficient fine-tuning of large-scale pre-trained language models. Nat. Machine Intelligence 5(3):220–235.Google Scholar
- (2024) Can LLM be a personalized judge? Preprint, submitted June 17, https://arxiv.org/abs/2406.11657.Google Scholar
- (2024) The Llama 3 herd of models. Preprint, submitted November 23, https://arxiv.org/abs/2407.21783.Google Scholar
- (2024) Determinants of LLM-assisted decision-making. Preprint, submitted February 27, https://arxiv.org/abs/2402.17385.Google Scholar
- (2025) Take caution in using LLMs as human surrogates. Proc. Natl. Acad. Sci. USA 122(24):e2501660122.Google Scholar
- Gemini Team, Anil R, Borgeaud S, Alayrac JB, Yu J, Soricut R, Schalkwyk J, (2023) Gemini: A family of highly capable multimodal models. Preprint, submitted December 19, https://arxiv.org/abs/2312.11805v1.Google Scholar
- Gemma Team, Kamath A, Ferret J, Pathak S, Vieillard N, Merhej R, Perrin S, (2025) Gemma 3 technical report. Preprint, submitted March 25, https://arxiv.org/abs/2503.19786.Google Scholar
- (2014) Generative adversarial nets. Adv. Neural Inform. Processing Systems 27:2672–2680.Google Scholar
- (2025) Designing LLM chains by adapting techniques from crowdsourcing workflows. ACM Trans. Comput.-Human Interaction 32(3):1–57.Google Scholar
- (2025) DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645(8081):633–638.Google Scholar
- (2020) Denoising diffusion probabilistic models. Adv. Neural Inform. Processing Systems 33:6840–6851.Google Scholar
- (2024) Bridging language and items for retrieval and recommendation. Preprint, submitted March 6, https://arxiv.org/abs/2403.03952.Google Scholar
- (2006) The rise of crowdsourcing. Wired Magazine 14(6):176–183.Google Scholar
- (2024) Quantifying the persona effect in LLM simulations. Proc. 62nd Annu. Meeting Assoc. Comput. Linguistics, vol. 1 (Association for Computational Linguistics, Kerrville, TX), 10289–10307.Google Scholar
- (2021) LoRA: Low-rank adaptation of large language models. Preprint, submitted October 16, https://arxiv.org/abs/2106.09685.Google Scholar
- (2024) Intuitive fine-tuning: Towards unifying SFT and RLHF into a single process. Preprint, submitted May 20, https://arxiv.org/abs/2405.11870v1.Google Scholar
- (2014) Cost of quality in crowdsourcing. Human Comput. 1(2):283–314.Google Scholar
- (2015) Recommendation systems: Principles, methods and evaluation. Egyptian Informatics J. 16(3):261–273.Google Scholar
- (2025) LLM economist: Large population models and mechanism design in multi-agent generative simulacra. NeurIPS 2025 Workshop Algorithmic Collective Action (San Diego).Google Scholar
- (2014) Adam: A method for stochastic optimization. Internat. Conf. Learn. Representations (San Diego).Google Scholar
- (2013) Auto-encoding variational Bayes. Preprint, submitted December 20, https://arxiv.org/abs/1312.6114v1.Google Scholar
- (2023) Understanding the effects of RLHF on LLM generalisation and diversity. Preprint, submitted October 10, https://arxiv.org/abs/2310.06452v1.Google Scholar
- (2021) Normalizing flows: An introduction and review of current methods. IEEE Trans. Pattern Anal. Machine Intelligence 43(11):3964–3979.Google Scholar
- (2023) Do LLM agents exhibit social behavior? Preprint, submitted December 23, https://arxiv.org/abs/2312.15198v1.Google Scholar
- (2024a) A comparative study on annotation quality of crowdsourcing and LLM via label aggregation. ICASSP 2024 IEEE Internat. Conf. Acoustics Speech Signal Processing (IEEE, Piscataway, NJ), 6525–6529.Google Scholar
- (2024b) Human-LLM hybrid text answer aggregation for crowd annotations. Preprint, submitted October 22, https://arxiv.org/abs/2410.17099.Google Scholar
- (2025) LLM generated persona is a promise with a catch. Preprint, submitted March 18, https://arxiv.org/abs/2503.16527.Google Scholar
- (2023) DPM-OT: A new diffusion probabilistic model based on optimal transport. Proc. IEEE/CVF Internat. Conf. Comput. Vision (IEEE, Piscataway, NJ), 22624–22633.Google Scholar
- (2024) Getting more juice out of the SFT data: Reward learning from human demonstration improves SFT for LLM alignment. Adv. Neural Inform. Processing Systems 37:124292–124318. Google Scholar
- (2025) Cultural bias in large language models: A comprehensive analysis and mitigation strategies. J. Transcultural Commun. 3(2):224–244.Google Scholar
- (2024a) DeLLMa: Decision making under uncertainty with large language models. Internat. Conf. Learn. Representations (Singapore).Google Scholar
- (2024b) CoachLM: Automatic instruction revisions improve the data quality in LLM instruction tuning. 2024 IEEE 40th Internat. Conf. Data Engrg. (IEEE, Piscataway, NJ), 5184–5197.Google Scholar
- (2024c) Make LLM a testing expert: Bringing human-like interaction to mobile GUI testing via functionality-aware decisions. ICSE’24 Proc. IEEE/ACM 46th Internat. Conf. Software Engrg. (Association for Computing Machinery, New York), 1–13.Google Scholar
- (2025) Beyond believability: Accurate human behavior simulation with fine-tuned LLMs. Preprint, submitted March 26, https://arxiv.org/abs/2503.20749v1.Google Scholar
- (2024) Improving linguistic diversity of large language models with possibility exploration fine-tuning. Preprint, submitted December 4, https://arxiv.org/abs/2412.03343.Google Scholar
- (2024) Generative AI voting: Fair collective choice is resilient to LLM biases and inconsistencies. Preprint, submitted May 31, https://arxiv.org/abs/2406.11871v1.Google Scholar
- (2015) Image-based recommendations on styles and substitutes. Proc. 38th Internat. ACM SIGIR Conf. Res. Development Inform. Retrieval (Association for Computing Machinery, New York), 43–52.Google Scholar
- (2024) AI emerges as the frontier in behavioral science. Proc. Natl. Acad. Sci. USA 121(10):e2401336121.Google Scholar
- (2024) LLMs to replace crowdsourcing for parallel data creation? The case of text detoxification. Findings Assoc. Comput. Linguistics EMNLP 2024 (Association for Computational Linguistics, Kerrville, TX), 14361–14373.Google Scholar
- (2024) Nomic embed: Training a reproducible long context text embedder. Preprint, submitted February 2, https://arxiv.org/abs/2402.01613v1.Google Scholar
- (2024) Evaluating persona prompting for question answering tasks. Internat. Conf. Artificial Intelligence Soft Comput.Google Scholar
- (2023) Does writing with language models reduce content diversity? Internat. Conf. Learn. Representations (Vienna, Austria).Google Scholar
- (2023) When do annotator demographics matter? Measuring the influence of annotator demographics with the POPQUORN dataset. Proc. 17th Linguistic Annotation Workshop (Association for Computational Linguistics, Kerrville, TX), 252–265.Google Scholar
- (1946) The elementary statistics of majority voting. J. Roy. Statist. Soc. 109(1):53–57.Google Scholar
- (2024) AI and the problem of knowledge collapse. Preprint, submitted April 22, https://arxiv.org/abs/2404.03502.Google Scholar
- (2025) AgentSociety: Large-scale simulation of LLM-driven generative agents advances understanding of human behaviors and society. Preprint, submitted February 12, https://arxiv.org/abs/2502.08691v1.Google Scholar
- (2024) An agentic AI-based multi-agent framework for recommender systems. 2024 IEEE Internat. Conf. Big Data (IEEE, Piscataway, NJ), 5375–5382.Google Scholar
- (2024) Direct preference optimization: Your language model is secretly a reward model. Adv. Neural Inform. Processing Systems 36:53728–53741.Google Scholar
- (2017) Proximal policy optimization algorithms. Preprint, submitted August 28, https://arxiv.org/abs/1707.06347.Google Scholar
- (2024) Curated LLM: Synergy of LLMs and data curation for tabular augmentation in low-data regimes. Proc. 41st Internat. Conf. Machine Learn. (PMLR, New York), 44060–44092.Google Scholar
- (2024) RAH! RecSys–Assistant–Human: A human-centered recommendation framework with LLM agents. IEEE Trans. Comput. Soc. Systems 11(5):6759–6770.Google Scholar
- (2025) Evaluating the diversity and quality of LLM generated content. Preprint, submitted April 16, https://arxiv.org/abs/2504.12522v1.Google Scholar
- (2019) Legal and ethical issues surrounding the use of crowdsourcing among healthcare providers. Health Informatics J. 25(4):1618–1630.Google Scholar
- (2023) LLM-planner: Few-shot grounded planning for embodied agents with large language models. Proc. IEEE/CVF Internat. Conf. Comput. Vision (IEEE, Piscataway, NJ), 2998–3009.Google Scholar
- (2024) Building better AI agents: A provocation on the utilisation of persona in LLM-based conversational agents. Proc. 6th ACM Conf. Conversational User Interfaces, vol. 35 (Association for Computing Machinery, New York), 1–6.Google Scholar
- (2024) Simulation-based exploration for aggregation algorithms in human+AI crowd: What factors should we consider for better results? AAAI HCOMP 12th AAAI Conf. Human Comput. Crowdsourcing 2024 (AAAI Press, Washington, DC).Google Scholar
- (2023) LLaMA: Open and efficient foundation language models. Preprint, submitted February 27, https://arxiv.org/abs/2302.13971.Google Scholar
- (2018) Making better use of the crowd: How crowdsourcing can advance machine learning research. J. Machine Learn. Res. 18(193):1–46.Google Scholar
- (2025) Prevalence and prevention of large language model use in crowd work. Commun. ACM 68(3):42–47.Google Scholar
- (2024) United in diversity? Contextual biases in LLM-based predictions of the 2024 European Parliament elections. Preprint, submitted August 29, https://arxiv.org/abs/2409.09045v1.Google Scholar
- (2025a) Multilingual prompting for improving LLM generation diversity. Preprint, submitted September 27, https://arxiv.org/abs/2505.15229.Google Scholar
- (2025b) YuLan-OneSim: Towards the next generation of social simulator with large language models. Preprint, submitted August 26, https://arxiv.org/abs/2505.07581.Google Scholar
- (2025c) Privacy risks of LLM-empowered recommender systems: An inversion attack perspective. Proc. 19th ACM Conf. Recommender Systems (Association for Computing Machinery, New York), 812–821.Google Scholar
- (2025d) User behavior simulation with large language model-based agents. ACM Trans. Inform. Systems 43(2):1–37.Google Scholar
- (2023) Self-consistency improves chain of thought reasoning in language models. Internat. Conf. Learn. Representations (Kigali, Rwanda).Google Scholar
- (2022) Chain-of-thought prompting elicits reasoning in large language models. Adv. Neural Inform. Processing Systems 35:24824–24837.Google Scholar
- (2009) Whose vote should count more: Optimal integration of labels from labelers of unknown expertise. Adv. Neural Inform. Processing Systems 22:2035–2043.Google Scholar
- (2023) A unified theory of diversity in ensemble learning. J. Machine Learn. Res. 24(359):1–49.Google Scholar
- (2025) LLM fine-tuning: Concepts, opportunities, and challenges. Big Data Cognitive Comput. 9(4):87.Google Scholar
- (2023) LLMs as workers in human-computational algorithms? replicating crowdsourcing pipelines with LLMs. Preprint, submitted July 19, https://arxiv.org/abs/2307.10168v1.Google Scholar
- (2020) Privacy in crowdsourcing: A review of the threats and challenges. Comput. Supported Cooperative Work 29:263–301.Google Scholar
- (2024) Understanding the performance and estimating the cost of LLM fine-tuning. Preprint, submitted August 8, https://arxiv.org/abs/2408.04693.Google Scholar
- (2023) The dark side of recruitment in crowdsourcing: Ethics and transparency in micro-task marketplaces. Comput. Supported Cooperative Work 32(3):439–474.Google Scholar
- (2024) Can large language model agents simulate human trust behavior? Adv. Neural Inform. Processing Systems 37:15674–15729.Google Scholar
- (2024a) On the role of large language models in crowdsourcing misinformation assessment. Proc. Internat. AAAI Conf. Web Soc. Media 18:1674–1686.Google Scholar
- (2024b) Echoes in AI: Quantifying lack of plot diversity in LLM outputs. Preprint, submitted December 31, https://arxiv.org/abs/2501.00273v1.Google Scholar
- (2024a) LLM voting: Human choices and AI collective decision-making. Proc. AAAI/ACM Conf. AI Ethics Society, vol. 7 (AAAI Press, Washington, DC), 1696–1708.Google Scholar
- (2024b) Designing digital voting systems for citizens: Achieving fairness and legitimacy in participatory budgeting. Digital Government Res. Practice 5(3):1–30.Google Scholar
- (2025) Qwen3 technical report. Preprint, submitted May 14, https://arxiv.org/abs/2505.09388.Google Scholar
- (2024) How reliable is human feedback for aligning large language models? Preprint, submitted October 2, https://arxiv.org/abs/2410.01957v1.Google Scholar
- (2018) Semi-implicit variational inference. Proc. 35th Internat. Conf. Machine Learn. (PMLR, New York), 5660–5669.Google Scholar
- (2024) LoFIT: Localized fine-tuning on LLM representations. Adv. Neural Inform. Processing Systems 37:9474–9506.Google Scholar
- (2023) Why Johnny can’t prompt: How non-AI experts try (and fail) to design LLM prompts. CHI’23 Proc. 2023 CHI Conf. Human Factors Comput. Systems (Association for Computing Machinery, New York), 1–21.Google Scholar
- (2024) Combining large language models and crowdsourcing for hybrid human-AI misinformation detection. Proc. 47th International ACM SIGIR Conf. Res. Development Inform. Retrieval (Association for Computing Machinery, New York), 2332–2336.Google Scholar
- (2024a) Neural prompt search. IEEE Trans. Pattern Anal. Machine Intelligence 47(7):5268–5280.Google Scholar
- (2014) Spectral methods meet EM: A provably optimal algorithm for crowdsourcing. J. Machine Learn. Res. 17(102):1–44.Google Scholar
- (2024b) On generative agents in recommendation. Proc. 47th Internat. ACM SIGIR Conf. Res. Development Inform. Retrieval (Association for Computing Machinery, New York), 1807–1817.Google Scholar
- (2025a) SocioVerse: A world model for social simulation powered by LLM agents and a pool of 10 million real-world users. Preprint, submitted July 15, https://arxiv.org/abs/2504.10157.Google Scholar
- (2025b) NoveltyBench: Evaluating creativity and diversity in language models. Preprint, submitted August 9, https://arxiv.org/abs/2504.05228.Google Scholar
- (2025c) LLM-powered user simulator for recommender system. Proc. AAAI Conf. Artificial Intelligence, vol. 39 (AAAI Press, Washington, DC), 13339–13347.Google Scholar
- (2024) An electoral approach to diversify LLM-based multi-agent collective decision-making. Preprint, submitted October 19, https://arxiv.org/abs/2410.15168.Google Scholar
- (2025) Pareto prompt optimization. Internat. Conf. Learn. Representations (Singapore).Google Scholar
- (2023) SOTOPIA: Interactive evaluation for social intelligence in language agents. Internat. Conf. Learn. Representations (Vienna, Austria).Google Scholar

