Artificial Intelligence-Supported Case Discussions: Harnessing Artificial Intelligence for Critical Engagement

Published Online:https://doi.org/10.1287/ited.2025.0177

Abstract

The case method has been successfully used in operations management courses and is central to business education, yet its effectiveness depends on consistent student preparation. Many students engage only superficially with cases, and the rise of generative artificial intelligence (AI) tools, such as ChatGPT, further complicates the challenge—providing shortcuts that yield summaries without fostering deeper understanding of trade-offs. This study investigates whether AI can instead be used to strengthen the case method as a pedagogical tool. Partnering with an AI-based upskilling platform, we conducted a field experiment using an AI-supported “quest” structure. We found that students achieved higher quiz scores when using the AI tool to support case preparation than when completing the same type of task without AI-supported preparation. For a group of 38 participants, on average, AI use was associated with an improvement of almost 1 point on a 10-point quiz. This difference was statistically significant using a paired t-test (t(37) = 4.50, p < 0.001). In addition, scores in the AI-supported condition were statistically less variable (standard deviation (SD) = 2.29 versus SD = 1.37), suggesting that AI support may help promote more consistent learning outcomes across students. These findings suggest that rather than undermining the case method, generative AI can enhance case learning when integrated deliberately into preclass preparation.

1. Introduction

Case-based teaching has long been a cornerstone of business education. Even operations research programs have successfully used case-based courses (Frances and Terekhov 2019). By placing students in the role of decision makers, cases provide an immersive opportunity to consider trade-offs, evaluate alternatives, and debate real-world business situations. When implemented effectively, case-based learning promotes not only content knowledge but also, critical thinking and communication skills essential for managerial and operations research practice.

Yet, the benefits of the case method are uneven (Booth et al. 2001). Students who only skim the material or arrive underprepared often contribute little to classroom discussion, lowering its value for themselves and their peers. Inadequate preparation can even be counterproductive as students come away with partial or distorted views of the problem. Inadequate preparation for an in-class discussion of a case is even more acute today. Many professors have noted that today’s students struggle to read long texts and that they also struggle with reading comprehension (Supiano 2024).

These same students have also never known a world without the internet or smart devices. As such, they have quickly adopted the use of generative artificial intelligence (AI) technologies, such as ChatGPT (Abd-Alrazaq et al. 2023). These tools are powerful summarizers and question-answering systems, and their ease of access makes it tempting for students to shortcut the preparation process—extracting key facts without engaging the deeper logic of the trade-offs (Cotton et al. 2024).

This tension raises important questions for business educators.

  • Is the case method still viable in an era when generative AI can instantly summarize any reading?

  • How can instructors preserve or even strengthen the developmental benefits of case preparation when students have access to AI shortcuts?

In this paper, we argue that generative AI need not undermine the case method. On the contrary, when carefully designed, AI can become a pedagogical partner that enhances preparation and deepens learning (Popenici and Kerr 2017, Zawacki-Richter et al. 2019, Crompton and Burke 2023). Working with an AI-based workforce training and productivity platform, we explored whether AI-supported preclass preparation could (1) improve comprehension of case content and (2) broaden participation across the classroom.

In this research, we explored the impact of an AI-supported “quest” structure that guided students to ask more relevant questions, uncover key concepts, and practice articulating answers before an in-class discussion. We use a randomized, controlled A/B test that compared students preparing with the AI tool against using traditional methods, measuring both quiz performance and in-class participation.

Our analysis yields several insights. First, we find that AI-supported preparation raised quiz scores by nearly a full point on average, an effect concentrated in the quiz for the case study that students found more difficult to answer. Second, quiz scores were also significantly less variable under AI-supported preparation. This points toward more consistent learning outcomes across students. Finally, evidence from one case discussion suggests that AI-supported preparation increases verbal engagement in class.

The remainder of the paper is organized as follows. We review the related literature in Section 2. Section 3 describes our experimental methodology, including the AI platform and study design. We present our results on quiz performance and in-class participation in Section 4. Section 5 discusses these findings and their implications for instructors. We conclude with directions for future research in Section 6.

2. Related Literature

Studies show case-based learning to be an effective method of learning (Paulus and Phipps 2008, Thistlethwaite et al. 2012). One benefit is the growth of critical thinking (Popil 2011, Herreid and Schiller 2013). Other studies counter that case-based learning has no benefits over traditional learning (Rhodes et al. 2020). More recently, Pilz et al. (2024) surveyed over 200 business school lecturers and concluded that contrary to the belief that cases are used primarily to develop social skills, their research found that business lecturers were using cases to impart content knowledge.

When commercial products, like ChatGPT, went viral in 2023, we saw the introduction of generative AI in the form of chatbots reach the masses (Dempere et al. 2023). Not long after, schools were banning the use of ChatGPT and other large language models by students (O’Brien 2024). Although readily available case solutions have historically been a problem for case-based learning, AI tools have made it easier for students to write a case analysis and submit it as their own (Lafkas 2024). This risk is distinct from the one that we study here; rather than allowing students to outsource case analysis to AI, the platform in our study is designed to structure how students engage with the case themselves.

Some researchers have explored the benefit of generative AI and specific tools, like ChatGPT (Kasneci et al. 2023, Tlili et al. 2023, Kosmyna et al. 2025, Wang and Fan 2025). They find that the benefits include the development of critical thinking and problem-solving skills. Both are skills needed to deeply discuss a business case in an in-class setting. Gerlich (2025) cautions, however, that frequent AI usage can lead to diminished critical thinking. Others have found evidence that the use of generative AI reduces grade dispersion in math courses (Bastani et al. 2025). This reduction in variation has also been seen in professional settings. Noy and Zhang (2023) found a reduction in the variation of time for midlevel professional writing tasks, and Fruits and Stout (2026) observed a reduction in the variation of customer call center issue resolution time with generative AI. This is the trade-off that our AI-supported quest structure is designed to navigate; rather than allowing passive summarization, the platform requires students to actively question and articulate answers before advancing, a design intended to build critical thinking rather than substitute for it. Others see potential for enhanced and customized learning with generative AI and chatbots. Chen et al. (2023) interviewed 215 undergrad students and discovered an openness to the use of chatbots as intelligent student assistants. Some instructors have shared how students can constructively use commercial large language models to help solve cases (Weinstein et al. 2025). However, most business cases in operations and supply chains are about critical thinking and rarely have one “right” answer. Instead, they are intended to generate deep learning during the in-class discussion. Along this line, others believe that the use of generative AI can allow students and instructors “to dig more deeply into the material” (Lafkas 2024) within the class time.

Akiba (2025) looked at chatbots to help class participation in an introductory psychology course of 25 students. He reached a mixed conclusion that AI-assisted participation interventions demonstrate both promise and limitations. However, this qualitative study only had four participants.

Most of the literature on case-based learning and AI looks at medical cases as opposed to business cases. Kononowicz et al. (2019), Plackett et al. (2022), and Wei et al. (2025) all found AI to be effective in increasing the clinical reasoning of students. Here, AI took the form of virtual patients. Higashitsuji et al. (2025) looked at ChatGPT for creating healthcare cases for nursing students. Although the sample size was very small, they found that the time to create a case was reduced without an impact on quality. Hassoulas et al. (2025) studied a variety of technologies, such as three-dimensional anatomy, AI chatbots, and Virtual Reality headsets, along with case-based learning and found that the students given this treatment performed better than the conventional group when it came to case-based learning.

Finally, Lang et al. (2024) and Jayasinghe et al. (2025) explored the use of AI to write teaching cases. These studies, however, focus primarily on medical education or on general attitudes toward AI chatbots rather than on how AI can structure case preparation itself.

3. Methodology

3.1. Pedagogical Setting

This study was conducted using a randomized crossover field experiment in an educational setting. The same experiment was run in the same graduate-level course (MGT 6772—Managing the Resources of the Technological Firm) with the same instructor for two consecutive semesters with a video recording of the second semester’s in-class discussion. Each participant read two business cases, a case about Dropbox (Case A) (Eisenmann 2012) and a case about NCR (Case B) (Collins et al. 2015). Participants were randomly assigned the AI platform’s quest structure for one case and no AI platform for the other case. Then, all participants took a quiz on both cases. This design enables within-student comparisons while counterbalancing participant differences and case difficulty. Finally, each case was discussed in class.

The graduate students in the class were randomly divided into two groups by the instructor. All of the students were instructed to read both Case A and Case B, and then, the students were directed to the AI platform. When the first group used the AI platform, they were only given access to the AI platform for Case A. When the second group used the AI platform, they were only given access to the AI platform for Case B.

At the start of class, all students in attendance took a quiz on Case A. Case A was then discussed. Following this, all students in attendance took a quiz on Case B, and it was discussed. In addition, the summer class was video recorded, and the recording was used to get verbal participation data.

We then employed a within-students analysis for quiz scores. We had expected to also use a within-students analysis for verbal participation, but technical issues prevented a full recording of the Case B discussion; therefore, we had to employ a between-students design for verbal participation.

3.2. How the AI Platform Works

This section describes the structure that students experienced when using the AI platform for either Case A or Case B.

  • Generating instructor-configured quests. Each platform is divided into three to five “quests.” Quests allow the student to explore deeper a key question from the case. For example, the platform for Case A yielded the following quest questions.

    • Why was Dropbox created, and what were the initial challenges?

    • What contributed to the success of Dropbox’s freemium model and referral program?

    • What distinguished Dropbox from competitors in the cloud storage market?

      The quests along with all of the content below are AI generated, but they are reviewed and tweaked by the instructor before assigning it to students.

  • Each quest starts with a self-rating and then, a podcast narrating a concept map. Within each quest, the student is first asked to self-rate how well they know the topic on a scale from one to five: one (novice), two (basic), three (intermediate), four (proficient), or five (master). Then, the student is given an AI-generated podcast (introduced midway through this study; see below) narrating an AI-generated concept map. For example, this is the start of the first quest above

  • Next, the student has to critically engage by asking questions until the concept map turns gold. The student is prompted to ask questions until all concepts are gold as seen in Figure 1. To help them ask questions, clicking on a gray concept reveals some question stems (see Figure 2), but the student is still responsible for asking a question. If their question is relevant, they get encouragement (such as “Right on point!”), and if the answer covered some hidden facts, the corresponding concepts turn gold. To gamify the experience, each revealed fact earns 10 coins. Coins can be used to buy question suggestions (50 coins per suggestion).

  • To finish the quest, the student has to give 30-word answers to three AI-generated questions. After the exploration phase, the quest enters the assessment stage. The student can take a quiz or play a visual game where they have to reassemble the concept map from memory. However, the mandatory assessment is to give open-ended answers (∼30 words) to AI-generated questions. The student gets instant feedback on what they said well and what they left out (Figure 3).

  • Postquest self-rating and app rating. After each quest, the student was asked to rate themselves now on the same scale from one (novice) to five (master). They are also asked if they would like more lessons to be quest based. The qualitative results from these ratings are not explored in this paper.

Figure 1. Sample Concept Map
Figure 2. Sample Question Stem
Figure 3. Sample AI Tool Assessment

3.3. Standards Compliance

The study was reviewed and approved under two institutional review board protocols (IRB2025-379 and IRB2025-430). This approval ensured that all data collection and analysis procedures adhered to ethical research standards in educational settings.

One of the authors is an employee of the AI upskilling platform used in this research. The other two authors are university faculty with no employment relationship outside their university.

4. Results

There were 68 students in total between both the spring and summer sections of MGT 6772. For the experiment, all of the students were randomly divided into Group A (34 students) and Group B (34 students). Of these two equal groups, 38 students consented to the experiment, used the AI tool, took both quizzes, and for summer, were in class for the discussion. These 38 students are referred to as the participants. Coincidentally, this provided an even split of 19 participants who used the AI platform with Case A and 19 participants who used the AI platform with Case B. The remaining 30 students in the courses did not consent, consented but did not use the AI tool, did not take both quizzes, or did not attend class on the day of the discussion.

In the summer MGT 6772 course, the in-class case discussions were also recorded, but the video recording unexpectedly cut off shortly into the Case B discussion. This meant that only Case A had the full recording of the in-class discussion. Because of this, we could not perform an analysis of the Case B in-class verbal discussion.

4.1. Quiz Performance

Quiz performance was measured using a 10-question multiple-choice quiz (maximum score = 10) administered prior to each case discussion. Because each participating student completed one case with AI support and one case without AI support, we focused on within-student comparisons to estimate the effect of AI use on quiz performance.

Across the 38 participating students, mean (M) quiz scores were higher when students used the AI tool (M = 8.92, standard deviation (SD) = 1.36) than when they did not use the AI tool (M = 7.84, SD = 2.26). The mean within-student improvement associated with AI use was 1.08 of 10 points (95% confidence interval = 0.59–1.56). This difference was statistically significant using a paired t-test (t(37) = 4.50, p < 0.001). Because the distribution of within-student differences did not follow a normal distribution, we also conducted a Wilcoxon signed-rank test, which yielded a consistent result (W = 31.0, p < 0.001).

A linear mixed-effects model with a random intercept for student was used to examine the effects of AI-assisted preparation, quiz version, and their interaction on quiz performance. Results indicated a significant main effect of the quiz (β = −1.68, standard error (SE) = 0.56, z = −2.99, p = 0.003), indicating that Quiz A was more difficult than Quiz B. Although the overall AI main effect was not significant (β = −0.37, SE = 0.56, z = −0.65, p = 0.513), a significant AI × quiz interaction emerged (β = 2.90, SE = 1.02, z = 2.84, p = 0.005). Follow-up interpretation of the interaction showed that AI improved performance on Quiz A by an estimated 2.53 points, representing an approximately 25% increase on a 10-point assessment, whereas the improvement was not statistically significant for Quiz B.

For Case A, the 19 students who used the AI tool achieved substantially higher quiz scores (M = 9.53, SD = 0.84) than the 19 students who did not use the AI tool for that case (M = 7.00, SD = 2.75, n = 19), corresponding to an average difference of 2.53 of 10 points. This difference was statistically significant using both a Welch two-sample t-test (p = 0.00095) and a Mann–Whitney U test (p = 0.00133).

In contrast, for Case B, average quiz scores were similar across conditions. Students who used the AI tool for Case B scored an average of 8.32 points (SD = 1.53) compared with 8.68 points (SD = 1.20) for students who did not use the AI tool, yielding a difference of −0.36 points that was not statistically significant (Welch t-test p = 0.42; Mann–Whitney U p = 0.54).

In addition to differences in mean performance, we also examined differences in the variability of quiz scores. Across the 38 participating students, quiz scores exhibited greater dispersion in the non-AI condition (SD = 2.29) than in the AI-supported condition (SD = 1.37). To assess statistical meaningfulness, we conducted the Levene test for equality of variances, which is robust to nonnormality. The test rejected the null hypothesis of equal variances (p = 0.047), indicating that quiz scores were significantly less variable when students used the AI tool. Figure 4 shows the variance graphically with a box plot.

Figure 4. Distribution of Quiz Scores with and Without AI Support

4.2. In-Class Participation (Summer 2025, Case A)

For the summer 2025 cohort, we analyzed verbal in-class participation during the Case A discussion. Verbal participation was measured by reviewing the class recording and counting each distinct verbal contribution made by each student during the discussion.

Because participation data were only collected for Case A, our analysis here uses between-student comparisons. A total of 16 participants were present. Eight participants were assigned support, and eight participants were assigned no AI. This even eight to eight split was a coincidence of attendance and not by design.

To better understand how the AI tool influenced verbal participation, verbal responses were examined across the extensive margin (whether a student participated at least once), the intensive margin (the number of responses among students who participated), and the overall margin (the number of responses across all students). Students assigned Case A were more likely to contribute to the discussion (87.5%, seven of eight students) than students not assigned the case (37.5%, three of eight students); however, this difference did not reach statistical significance according to the Fisher exact test (p = 0.119). Among students who participated, those assigned Case A made more discussion contributions on average (3.71 versus 2.33 responses, respectively), although the difference was also not statistically significant based on a Mann–Whitney U test (U = 3.5, p = 0.100). However, when all students were considered, including those who did not participate, students assigned AI for Case A contributed significantly more discussion responses overall (3.25 versus 0.88 responses per student, respectively, Mann–Whitney U = 11.5, p = 0.012). Collectively, these findings suggest that the intervention primarily increased classroom engagement by encouraging more students to participate in the discussion (extensive margin), with a smaller increase in the number of contributions made by students who chose to speak (intensive margin), resulting in a statistically significant increase in overall classroom verbal participation.

5. Discussion

This study examined the impact of AI quest structured case preparation on student learning and engagement in a live classroom setting. Using a randomized crossover field experiment, we found that students scored approximately 1 point higher on a 10-point quiz when using the AI tool compared with when they did not. Because each participant prepared for one case with AI support and prepared for one case without AI support, this improvement reflects within-student gains rather than differences in baseline ability, strengthening the causal interpretation of the performance effect.

Beyond improvements in average performance, the AI quest structure was also associated with a reduction in performance variability. Quiz scores when AI preparation was used were less dispersed than scores in the non-AI condition, suggesting that AI-supported preparation may function as a form of instructional scaffolding that helps weaker-performing students close performance gaps. From an educational perspective, this finding is particularly interesting as it implies that AI tools may support more equitable learning outcomes rather than benefiting only a subset of high-performing students.

Although we focused primarily on within-participant effects, it is noted that the analysis indicated a different AI effect between the two cases. Quiz B produced significantly higher scores overall, indicating that it was likely a less difficult assessment than Quiz A. Consequently, many students were already scoring near the maximum possible score on the 10-point quiz, leaving limited opportunity for additional improvement. This ceiling effect may partially explain why AI-assisted preparation produced a substantial improvement on Quiz A but not on Quiz B. These findings suggest that AI may provide the greatest educational benefit when students engage with more cognitively demanding material rather than assessments on which baseline performance is already high. Although the study was not designed to isolate the mechanisms driving these differences, the results highlight the importance of aligning AI use with pedagogical goals and case design.

Lastly, data from the summer 2025 class provide insight into how AI quest preparation may influence student engagement during case discussions. Although students who used the AI tool verbally participated at a much higher rate and more often than those who did not use the AI tool, neither were statistically significant. However, there was statistical significance when looking at the overall verbal participation counts between those who used the tool and those who did not. This mixed result points toward a benefit to verbal engagement.

Taken together, the findings suggest that AI-supported case preparation may enhance learning along multiple dimensions: by increasing average performance, reducing variability in outcomes, and fostering greater in-class engagement. Importantly, the engagement results indicate that AI support may lower barriers to participation for some students, even though not all AI-supported students chose to speak. This nuanced pattern underscores that AI tools do not automatically produce engagement but may create conditions that make participation more likely.

5.1. Implications for Instructors

The results of this study suggest two practical implications for instructors considering the integration of AI tools into case-based pedagogy. First, an AI-supported quest structure for case preparation can enhance learning outcomes. Students using the AI tool achieved higher quiz scores and exhibited less variability in performance, indicating that AI tools may help raise the floor and level the playing field by supporting students who might otherwise struggle with case analysis.

Second, AI-supported preparation appears to influence how students engage in class discussions. In the observed in-class case discussion, among all students who made a verbal contribution, students assigned to AI-supported preparation contributed more frequently than those who did not use the tool. This suggests that this AI tool may increase students’ readiness or confidence to engage verbally, particularly in complex or open-ended cases. Instructors seeking to broaden participation may, therefore, find AI-supported case preparation to be a useful complement to discussion-based learning.

6. Conclusion and Future Directions

This study examined the effects of AI-supported case preparation on student learning and engagement in an authentic classroom setting. Using a randomized crossover field experiment, we found that students achieved higher quiz scores when using an AI tool to support case preparation than when completing the same type of task without AI-supported preparation. On average, AI use was associated with an improvement of approximately 1 point on a 10-point quiz, and scores in the AI-supported condition were less variable, suggesting that AI support may help promote more consistent learning outcomes across students.

Evidence from in-class participation during one case discussion also suggests that AI-supported preparation is associated with greater participation intensity, although this engagement-related finding is limited in scope.

Several limitations of this study should be acknowledged. First, the study was conducted within a single course context and involved a limited number of cases, which may constrain the generalizability of the findings. Second, in-class verbal participation data were available for only one case in one semester, and the sample size was modest. Finally, the study focused on short-term quiz performance; longer-term learning outcomes and retention were not assessed.

Subsequent studies could examine AI-supported case discussions across multiple courses and disciplines. In addition, richer measures of engagement, including written contributions, and qualitative indicators of critical thinking could help clarify the mechanisms through which AI-supported quest preparation influences learning. Finally, research that examines longer-term outcomes, such as retention of concepts, would provide valuable insight into the educational impact of AI tools.

Overall, this study contributes to the growing literature on AI in education by demonstrating that AI-supported case preparation can enhance learning outcomes. When thoughtfully integrated into case-based pedagogy, AI tools appear to function not as substitutes for student learning but as a support that can foster preparation, participation, and more equitable learning outcomes.

References

  • Abd-Alrazaq A, AlSaad R, Alhuwail D, Ahmed A, Healy PM, Latifi S, Aziz S, Damseh R, Alabed Alrazak S, Sheikh J (2023) Large language models in medical education: Opportunities, challenges, and future directions. JMIR Medical Ed. 9(1):e48291.Crossref, Google Scholar
  • Akiba D (2025) ChatGPT told me to say it: AI chatbots and class participation apprehension in university students. Ed. Sci. 15(7):897.Google Scholar
  • Bastani H, Bastani O, Sungu A, Ge H, Kabakcı Ö, Mariman R (2025) Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proc. Natl. Acad. Sci. USA 122(26):e2422633122.Crossref, Google Scholar
  • Booth C, Bowie S, Jordan J, Rippin A (2001) The Use of the Case Method in Large and Diverse Undergraduate Business Programmes: Problems and Issues (A Report to the European Case Clearing House and the Foundation for Management Education) (University of West England Bristol Business School, Bristol, UK).Google Scholar
  • Chen Y, Jensen S, Albert LJ, Gupta S, Lee T (2023) Artificial intelligence (AI) student assistants in the classroom: Designing chatbots to support student success. Inform. Systems Frontiers 25(1):161–182.Crossref, Google Scholar
  • Collins DJ, Saden R, Shaffer M (2015) The Transformation of NCR (Harvard Business School, Boston).Google Scholar
  • Cotton DRE, Cotton PA, Shipway JR (2024) Chatting and cheating: Ensuring academic integrity in the era of ChatGPT. Innovations Ed. Teaching Internat. 61(2):228–239.Crossref, Google Scholar
  • Crompton H, Burke D (2023) Artificial intelligence in higher education: The state of the field. Internat. J. Educational Tech. High Ed. 20(1):22.Crossref, Google Scholar
  • Dempere J, Modugu K, Hesham A, Ramasamy LK (2023) The impact of ChatGPT on higher education. Frontiers Ed. 8:1206936.Crossref, Google Scholar
  • Eisenmann TR (2012) Dropbox: “It Just Works” (Harvard Business School, Boston).Google Scholar
  • Frances DM, Terekhov D (2019) A case-based undergraduate operations research course. INFORMS Trans. Ed. 19(2):67–80.Link, Google Scholar
  • Fruits E, Stout K (2026) AI, productivity, and labor markets: A review of the empirical evidence. Preprint, submitted March 16, https://doi.org/10.2139/ssrn.6323960.Google Scholar
  • Gerlich M (2025) AI tools in society: Impacts on cognitive offloading and the future of critical thinking. Societies 15(1):6.Crossref, Google Scholar
  • Hassoulas A, Crawford O, Hemrom S, de Almeida A, Coffey MJ, Hodgson M, Leveridge B, et al. (2025) A pilot study investigating the efficacy of technology enhanced case based learning (CBL) in small group teaching. Sci. Rep. 15(1):15604.Crossref, Google Scholar
  • Herreid CF, Schiller NA (2013) Case studies and the flipped classroom. J. College Sci. Teaching 42(5):62–66.Google Scholar
  • Higashitsuji A, Otsuka T, Watanabe K (2025) Impact of ChatGPT on case creation efficiency and learning quality in case-based learning for undergraduate nursing students. Teaching Learn. Nursing 20(1):e159–e166.Crossref, Google Scholar
  • Jayasinghe S, Arm K, Gamage KAA (2025) Designing culturally inclusive case studies with generative AI: Strategies and considerations. Ed. Sci. 15(6):645.Google Scholar
  • Kasneci E, Sessler K, Küchemann S, Bannert M, Dementieva D, Fischer F, Gasser U, et al. (2023) ChatGPT for good? On opportunities and challenges of large language models for education. Learn. Individual Differences 103:102274.Crossref, Google Scholar
  • Kononowicz AA, Woodham LA, Edelbring S, Stathakarou N, Davies D, Saxena N, Tudor Car L, Carlstedt-Duke J, Car J, Zary N (2019) Virtual patient simulations in health professions education: Systematic review and meta-analysis by the digital health education collaboration. J. Medical Internet Res. 21(7):e14676.Crossref, Google Scholar
  • Kosmyna N, Hauptmann E, Yuan YT, Situ J, Liao X-H, Beresnitzky AV, Braunstein I, Maes P (2025) Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing task. Preprint, submitted June 10, https://doi.org/10.48550/arXiv.2506.08872.Google Scholar
  • Lafkas J (2024) Students are using Gen AI to prep cases. Don’t worry. Harvard Bus. Impact (August 29), https://www.hbsp.harvard.edu/inspiring-minds/students-are-using-gen-ai-to-prep-cases-dont-worry.Google Scholar
  • Lang G, Triantoro T, Sharp JH (2024) Large language models as AI-powered educational assistants: Comparing GPT-4 and Gemini for writing teaching cases. J. Inform. Systems Ed. 35(3):390–407.Crossref, Google Scholar
  • Noy S, Zhang W (2023) Experimental evidence on the productivity effects of generative artificial intelligence. Science 381(6654):187–192.Crossref, Google Scholar
  • O’Brien M (2024) 2023: The year we played with artificial intelligence—And weren’t sure what to do about it. AP News (December 6), https://apnews.com/article/ai-2023-artificial-intelligence-chatgpt-dangers-565ff5b817b5db0d4e74829ae3d68611.Google Scholar
  • Paulus T, Phipps G (2008) Approaches to case analyses in synchronous and asynchronous environments. J. Comput.-Mediated Comm. 13(2):459–484.Crossref, Google Scholar
  • Pilz M, Tögel J, Albers S, van den Oord S, Cramer T, Vítečková K (2024) Teaching with business cases in higher education: Expectations and practical implementation by lecturers of management. Internat. J. Management Ed. 22(3):101068.Google Scholar
  • Plackett R, Kassianos AP, Mylan S, Kambouri M, Raine R, Sheringham J (2022) The effectiveness of using virtual patient educational tools to improve medical students’ clinical reasoning skills: A systematic review. BMC Medical Ed. 22(1):365.Crossref, Google Scholar
  • Popenici SAD, Kerr S (2017) Exploring the impact of artificial intelligence on teaching and learning in higher education. Res. Practice Tech. Enhanced Learn. 12(1):22.Crossref, Google Scholar
  • Popil I (2011) Promotion of critical thinking by using case studies as teaching method. Nurse Ed. Today 31(2):204–207.Crossref, Google Scholar
  • Rhodes A, Wilson A, Rozell T (2020) Value of case-based learning within STEM courses: Is it the method or is it the student? CBE Life Sci. Ed. 19(3):ar44.Crossref, Google Scholar
  • Supiano B (2024) Some assembly still required: How K-12 reforms and recent disruptions created Gen Z’s baffling habits. Chronicle Higher Ed. (December 20), https://www.chronicle.com/article/some-assembly-still-required.Google Scholar
  • Thistlethwaite JE, Davies D, Ekeocha S, Kidd JM, MacDougall C, Matthews P, Purkis J, Clay D (2012) The effectiveness of case-based learning in health professional education. A BEME systematic review: BEME Guide No. 23. Medical Teacher 34(6):e421–e444.Crossref, Google Scholar
  • Tlili A, Shehata B, Adarkwah MA, Bozkurt A, Hickey DT, Huang R, Agyemang B (2023) What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in education. Smart Learn. Environ. 10(1):15.Crossref, Google Scholar
  • Wang J, Fan W (2025) Retracted article: The effect of ChatGPT on students’ learning performance, learning perception, and higher-order thinking: Insights from a meta-analysis. Humanities Soc. Sci. Comm. 12(1):621.Crossref, Google Scholar
  • Wei H, Dai Y, Yuan K, Li KY, Hung KF, Hu EM, Lee AHC, Chang JWW, Zhang C, Li X (2025) AI-powered problem- and case-based learning in medical and dental education: A systematic review and meta-analysis. Internat. Dental J. 75(4):100858.Crossref, Google Scholar
  • Weinstein A, Brotspies HV, Gironda JT (2025) Do your students know how to analyze a case with AI—And still learn the right skills? A framework for using Gen AI to support, not replace, students’ critical thinking. Harvard Bus. Impact (April 14), https://www.hbsp.harvard.edu/inspiring-minds/framework-analyze-cases-using-ai-enhance-decision-making-skills.Google Scholar
  • Zawacki-Richter O, Marín VI, Bond M, Gouverneur F (2019) Systematic review of research on artificial intelligence applications in higher education—Where are the educators? Internat. J. Educational Tech. High Ed. 16(1):39.Crossref, Google Scholar