Assessing the Quality of Large Language Models and Human Inputs to a Decision: A Proposed Framework and Two Benchmark Case Studies
References
- (2013) Teaching decision-making with social networks. ORMS Today 40(6):26–29.Google Scholar
- (2024) Large language models in healthcare and medical domain: A review. Informatics 11(3):57.Crossref, Google Scholar
- (2023) Man ends his life after an AI chatbot ‘encouraged’ him to sacrifice himself to stop climate change. EuroNews (March 31), https://www.euronews.com/next/2023/03/31/man-ends-his-life-after-an-ai-chatbot-encouraged-him-to-sacrifice-himself-to-stop-climate-.Google Scholar
- (2023) Lawyer used ChatGPT in court—And cited fake cases. A judge is considering sanctions. Forbes (June 8), https://www.forbes.com/sites/mollybohannon/2023/06/08/lawyer-used-chatgpt-in-court-and-cited-fake-cases-a-judge-is-considering-sanctions/.Google Scholar
- (2008) Generating objectives: Can decision makers articulate what they want? Management Sci. 54(1):56–70.Link, Google Scholar
- (2010) Improving the generation of decision objectives. Decision Anal. 7(3):238–255.Link, Google Scholar
- (1997) The effects of elicitation aids, knowledge, and problem content on option quantity and quality. Organ. Behav. Human Decision Processes 72(2):184–202.Crossref, Google Scholar
- (2024) Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nature Medicine 30(September):2613–2622.Crossref, Google Scholar
- (1988) Decision analysis: Practice and promise. Management Sci. 34(6):679–695.Link, Google Scholar
- (2016) Foundations of Decision Analysis (Pearson Education, Boston).Google Scholar
- (2007) Developing objectives and attributes. Edwards W, Miles RF, von Winterfeldt D, eds. Advances in Decision Analysis: From Foundations to Applications (Cambridge University Press, Cambridge, UK), 104–128.Crossref, Google Scholar
- (1988) Decision problem structuring: Generating options. IEEE Trans. Systems Man Cybernetics 18(5):715–728.Crossref, Google Scholar
- (2025) Tasks and roles in legal AI: Data curation, annotation, and verification. Preprint, submitted April 2, https://doi.org/10.48550/arXiv.2504.01349.Google Scholar
- (2024) AI is creating fake legal cases and making its way into real courtrooms, with disastrous results. Conversation (March 12), https://theconversation.com/ai-is-creating-fake-legal-cases-and-making-its-way-into-real-courtrooms-with-disastrous-results-225080.Google Scholar
- (2023) Large language models in finance: A survey. Wang G, Cucuringu M, Kurshan E, eds. Proc. Fourth ACM Internat. Conf. AI Finance (ACM, New York), 374–382.Google Scholar
- (2021) Ideation in the digital age: Literature review and integrative model for electronic brainstorming. Rev. Management Sci. 15(6):1431–1464.Google Scholar
- (2024) Can A.I. be blamed for a teen’s suicide? New York Times (October 23), https://www.nytimes.com/2024/10/23/technology/characterai-lawsuit-teen-suicide.html.Google Scholar
- (2016) Decision Quality (John Wiley, Hoboken, NJ).Google Scholar
- (2023) BloombergGPT: A large language model for finance. Preprint, submitted March 30, https://doi.org/10.48550/arXiv.2303.17564.Google Scholar
- (2024) How well do LLMs cite relevant medical references? An evaluation framework and analyses. Preprint, submitted February 3, https://doi.org/10.48550/arXiv.2402.02008.Google Scholar

