Practical Guidance for Writing Multiple-Choice Test Questions in Introductory Analytics Courses

Published Online:https://doi.org/10.1287/ited.2022.0274

Abstract

Test writing is a fundamental component of teaching. With increasing pressure to teach larger groups of students, conduct formal assessment of learning outcomes, and offer online and hybrid classes, there is a need for alternatives to constructed response problem-solving test questions. We believe that appropriate use of multiple-choice (MC) questions can alleviate some of these pressures. However, the results will only be acceptable to dedicated faculty if these questions are well written and effective in measuring student learning. In this article, we propose a structured framework consisting of six design principles that serve as a guide for writing MC items that promote higher-order thinking through the use of scenario-based and multistep problems. Then we demonstrate how to apply these principles to develop high-quality test questions in MC format for introductory analytics courses in statistics or operations research/management science.

1. Introduction

The “conventional” multiple-choice (MC) format (Haladyna and Rodriguez 2013) consists of a “stem” containing a question or partial sentence and a list of choices or “options,” only one of which is an appropriate response to the stem. The incorrect options are called “distractors.” MC test questions have a long history, originating with the application of Frederick Taylor’s techniques for engineering management to education. Prior to the start of World War I, essay exams were the primary mode of testing. However, essay questions came under increasing criticism for being inefficient and unreliable. In addition, educational researchers were advocating the development of norm-referenced, standardized tests for categorizing students (Madaus and O’Dwyer 1999). Since then, MC tests have become popular for assessment at all levels of education and for aptitude testing, licensing, and certification exams such as the INFORMS Certified Analytics Professional (CAP) programs.

The widespread adoption of MC tests arises from a number of advantages that they provide. They can be used to administer examinations to large groups at the same time with low-cost, computer-based scoring and immediate feedback. These attributes have become major advantages in higher education with the advent of formal assessment for accreditation, larger class sizes, and more distance education, particularly as a result of the continuing COVID-19 pandemic. MC questions (MCQs) also reduce the time required for testing (i.e., selecting an answer versus writing a response), thereby allowing more content to be covered and increasing the validity of assessment. Their results can be analyzed with statistical measures for consistency and reliability and used to guide curriculum and pedagogical changes for continuous improvement (Palocsay et al. 2020). They have also been shown to be effective in promoting student learning as formative assessment tools (Butler 2018). The test banks that accompany textbooks are a commonly used source of MC test questions for many instructors.

However, there are several issues associated with these test banks. They generally contain a limited number of questions, whereas online testing requires a large pool for randomization to minimize academic dishonesty (Miltenburg 2019). Moreover, MC items in published test banks are often criticized for emphasizing rote memorization (Simkin and Kuechler 2005), corresponding to the lowest cognitive level of the taxonomy of Bloom (1956) (i.e., remember). It is possible to construct MCQs that evaluate higher-level cognitive processes (Scully 2017), but it is difficult to write MC test questions of high quality that assess problem-solving abilities in analytics.

In this paper, we seek to address these challenges by providing a series of practical recommendations for effective development of MC test questions for assessment in introductory analytics classes. First, we review relevant research from the educational literature and discuss the use of MCQs for assessment testing. We then present and examine a set of design principles useful in creating effective MCQs for introductory analytics classes. We demonstrate how to apply these principles to write nontrivial MCQs testing a broad spectrum of cognitive skills in several areas of business analytics: probability, statistics, queuing theory, linear programming (LP), decision analysis, and simulation. Specific attention is given to showing how such questions can test problem-solving competencies. Finally, we briefly summarize issues associated with testing using MCQs and outline directions for future research on the comparative ability of MCQs to measure performance.

2. Assessment Testing with MCQs

Many educational researchers advocate using a systematic, three-stage approach for effective course planning referred to as “backward design” (McTighe and Wiggins 2012). Starting with deliberation on overall goals, instructors are asked to first establish learning objectives (LOs) that articulate the knowledge and skills students should have on completion of the course. Gronlund (2004) prescribes identification of course instructional objectives as the foundation of effective planning, teaching, and assessment. In particular, LOs, clearly stated in performance terms, are essential in developing assessment instruments, because they clarify standards for evaluating student achievement. Criteria for well-written LOs are often summarized by the SMART acronym taken from business strategy for setting management goals: specific, measurable, achievable, relevant, and timely (see Lawlor and Hornyak 2012 for historical background). LOs can also be characterized by their cognitive level in the taxonomy of Bloom (1956), as revised by Anderson and Krathwohl (2001). This framework consists of a hierarchy of six categories: remember, understand, apply, analyze, evaluate, and create.

After identifying desired learning outcomes, instructors must then address the question of how to assess student performance on the LOs. In the context of assessment, Anderson and Krathwohl (2001, p. 63) state that “two of the most important educational goals are to promote retention and to promote transfer (which when it occurs, indicates meaningful learning).” They go on to lament an overemphasis on retention and describe the difficulties associated with assessing LOs intended to advance students’ ability to “use what was learned to solve new problems” (p. 63). Researchers have continued to find evidence that validates this criticism in the sciences (Momsen et al. 2010, Lemons and Lemons 2013, Vanderbilt et al. 2013) and engineering (Swart 2010, Meda and Swart 2018).

Swart (2010) classified between 33% and 52% of final exam questions in a sequence of four electronics courses as targeting knowledge and comprehension, that is, the two lowest levels of Bloom’s hierarchy. Both course goals and test questions in introductory undergraduate biology classes were also categorized by level (Momsen et al. 2010). These authors rated 93% of assessment items and 69% of goals at the remember and understand levels. Furthermore, their analysis showed no statistically significant correlation between mean Bloom’s levels of course goals and assessments, even though faculty in the study had training on the backward design methodology. They identify the root cause of this disconnect as the pressure to use machine-graded tests, because of large class sizes, combined with the challenge of writing high-level MCQs.

More recently, Scully (2017) provides justification for MC items that are capable of assessing up to at least through the analyze (fourth) level of the revised Bloom’s taxonomy. He acknowledges that evaluation of advanced cognitive processes is domain dependent but recommends several generic strategies for producing MCQs that attain these higher levels: use of verbs associated with higher cognitive levels, flipping the stem and choices for recall items (i.e., presentation of specific instances in the stem with choices that require identification of an underlying concept), development of high-quality distractors, and transformation of single-neuron items to at least the two-neuron level. His advice informed our efforts to address some of the limitations of MCQs for assessment of quantitative problem-solving skills.

Recognition of the importance of developing well-constructed MCQs has generated extensive literature on the topic. Empirical research has investigated a range of topics that affect the quality of MCQs: the optimal number of answers (Rodriguez 2005), ordering of answer options (Wang 2019), effects of using flawed test items (Downing 2005), and statistical discrimination power of distractors (Gierl et al. 2017, Haladyna and Rodriguez 2021). It has also resulted in the generation of a host of recommendations and suggestions for writing these types of questions. There is general consensus supporting 45 guidelines originally from Haladyna and Downing (1989), revised (Haladyna et al. 2002, Haladyna and Rodriguez 2013), and recently condensed by Rodriguez and Albano (2017) into a list of 22 rules. These rules are organized into four categories: content, formatting and style, stem writing, and writing of choices. They are mostly common sense and aimed at preventing common flaws such as repeating word phrases in answer options, writing test items that are not independent, and inadvertently providing hints to correct answers. We have absorbed these guidelines into our proposed MCQ design principles.

Historically, there has been a great deal of attention paid to the process of writing effective MCQs in the medical sciences, with related publications dating back to the late 1970s (see references in Collins 2006). This interest is motivated by the heavy use of MC tests for assessment in graduate and continuing medical education (Ghidinelli et al. 2021). With respect to the potential limitations of MC testing, Morrison and Free (2001) review four criteria for fostering critical thinking in the nursing discipline using MCQs: focus test items on higher-order Bloom’s levels, require multilogical thinking, present plausible alternatives as distractors, and provide rationale for correct and incorrect answers (after testing).

The notion of multilogical thinking refers to the need to relate and apply more than one fact and/or concept to a situation. Burns (2010) uses similar terminology for classifying questions in anatomy using a metaphor based on the number of neurons that must be triggered to answer an MCQ. He states that “multiple-neuron questions encourage student use of higher order learning behaviors requiring some level of summarizing, synthesizing, and inferential reasoning” (Burns 2010, p. 332). We adopt this recommendation to require application of multiple concepts in an MCQ, referring to it as the level of ‘stacking’ in our design principles.

Criticism of MC testing is primarily based on an assumption that only constructed response (CR)-type questions can measure higher-order skills as defined by Bloom’s framework. Empirical research on the relationship between performance on MCQs and performance on CR questions has been inconclusive about their equivalency (Kuechler and Simkin 2010). However, a meta-analysis by Rodriguez (2003) identified a trend in association studies with higher correlations between results using the two formats as the similarity of items increased.

A typical CR question in analytics requires students to draw on a variety of mental faculties and integrate them. Accomplishing this level of testing with MCQs is not practical in some instances, such as for theoretical topics requiring mathematical proofs and derivations. But the situation is different in lower-level analytics courses where the final result of a CR question is a definitive outcome: the computed value of a quantity, the identification of a characteristic or term, or the association/comparison/contrast of specified ideas or quantities. Although some of these results may be obtained via a single cognitive step, most involve multiple steps performed in sequence with supporting work being shown. A single-step question is relatively easy to recast as an MCQ, but the more steps a problem involves, the more difficult that recasting process becomes because partial credit is hard to assign with an MCQ. Because the intermediate steps are not shown, even a small error somewhere in the work sequence can lead to an answer that essentially appears to be a random guess. This disproportionate penalty for minor errors is likely one reason that so many instructors either write single-step MCQs or do not use MCQs at all.

We approach the development of higher-order assessment testing in an MCQ environment by recognizing that the comparison of a CR question to a single MCQ is often not appropriate. Any multistep CR question has an implicit prompt—“what next?”—within it. To reframe it using MC methodology, the key idea is to create a description of a situation followed by a collection of MCQs based on it. Haladyna and Rodriguez (2013) refer to this collection as a “testlet” and state its primary advantage is its “capacity to model the kind of complex thinking found in a constructed-response subjectively-scored item” (p. 77). Writing questions with “vignettes” is also recommended for medical exams to test both recall and application of foundational and clinical knowledge (Paniagua and Swygert 2016, p. 33). We prefer the term scenario for such a collection, and in this article, we include not only sets in which each new MCQ continues the work begun by its predecessor but also sets where different MCQs use the same situation description to drill down to ask independent questions targeted at specific concepts or skills. One could argue that making the “what next?” prompt explicit in a scenario-based MCQ set provides too much guidance on how work on the problem should proceed. There is indeed a tradeoff here, and as we discuss, careful scenario design is needed to avoid oversimplification.

In the next section, we offer six design principles, shown in Figure 1, that provide guidance on developing MCQs based on LOs to effectively measure student learning in introductory analytics courses using both scenarios and stand-alone items. We incorporate the standard guidelines from Rodriguez and Albano (2017) into our principles, as indicated in Figure 1, and give examples of both well- and poorly constructed MCQs.

Figure 1. Design Principles for Effective MCQs

3. Six Design Principles for MCQ Writing in Analytics

3.1. Choose a Question Type Appropriate to the LO to Be Tested

One needs to have appropriately stated LOs objectives before writing an examination. If one can’t articulate what students are supposed to learn, how can one create questions that assesses the extent to which they have learned it? As noted earlier, it is important for an objective which a question is intended to test to be both specific and measurable. An objective such as “students will be able to understand and use probability” is neither specific nor measurable. Probability is a deep and wide field; we must specify what parts of it are to be mastered. And how does one measure if a student “understands probability”? We need to operationalize this objective by identifying what the student is supposed to be able to do. Example entries might include the following:

Students will be able to

  • (1) Determine whether events are complementary.

  • (2) Identify marginal, joint, and conditional probability statements in word problems and represent these probabilities using appropriate notation.

  • (3) Apply basic rules of probability to compute marginal, joint, and conditional probabilities.

After choosing the target LO, consider what type or level of complexity you wish the question to have. Although Bloom’s taxonomy is frequently referenced in this regard, it is limited in characterizing some test questions (Lemons and Lemons 2013). It suggests that the task of identify is at a lower level in the hierarchy than the task of calculate; however, identifying an appropriate queuing model, for example, is a substantially more advanced task than calculating the value of L = 1/(μ – λ) given values of μ and λ. For writing questions in basic analytics, we have found it more useful to characterize questions as shown in Table 1: reading-reasoning, memorization, identification, calculation, method-application pairing, and applied concept.

Table

Table 1. Types of MC Questions

Table 1. Types of MC Questions

Question typeDescriptionCommentsMinimal stacking example (partial, with answer)
Reading-reasoningBasic reading and reasoning skills, not course content per se.Provides easy entry to test, gets students thinking along the desired line, flags students with reading, reasoning or comprehension problems. All word problems include this in their stacks. Questions involving only this level are rare.We survey families that have more than two cars. What is the smallest number of cars such a family might have?
MemorizationTerminology, notation, definitionsLittle or no stacking. Flags students who have not done minimal test preparation.A solution that satisfies all of a program’s constraints is called (feasible.)
IdentificationLinking a term or concept to a given exampleInvolves minimal stacking with reading-reasoning and possibly with memorization. The foundation for a method-application pairing problem.The requirement that the respondent has at least three cars is an example of a (constraint.)
CalculationUse familiar formulas or mechanical procedures in straightforward waysAssesses the ability to follow a familiar “recipe”. Possibly involves competency with arithmetic, algebra, calculators, or spreadsheets. Stacking with lower levels or other calculations is variable and will increase difficulty while reducing the diagnostic value of a wrong answer.If x + y = 10 and 2x + 3y = 26, find y. (6)
Method-application pairingDetermine which tool, technique or concept is (or is not) useful for a given taskAssesses a practical understanding of the utility and limitations of a tool, technique or concept. Builds on both reading-reasoning and identification.You are asked to use LP to solve MAXIMIZE (x + 2)(x − 2) subject to 1 ≤ x ≤ 3. Your reply should be (“I can’t. The objective function is nonlinear.”)
Applied conceptApply a concept, tool, or calculation appropriate to the problemThe workhorse of a business analytics test, this problem type assesses practical competence. Builds on a method-application pairing. Stacking increases if the method-application pairing needed is left unstated, or if memorization or calculation skills are included in obtaining the correct answer. Careful selection of stacking level is essential.An LP with the objective of MAXIMIZE P = 4x + 7y has an optimal solution at x = 4, y = 10. We add to this program the constraint x ≥ 2. What is the effect on P? (no change)

In describing these question types, the Purpose/comments column of Table 1 refers to the notion of stacking, related to the idea of Morrison and Free (2001) of “multilogical thinking” and the number of neurons of Burns (2010). More coverage of stacking in MCQs is provided in Section 3.6, but in brief, higher levels of stacking correspond to more complexity and occur when a question requires the application of more concepts, skills, decisions, or steps, even if those steps are repetitive. The examples in Table 1 are from linear programming (LP) and were chosen to clarify descriptions of the question types, keeping the degree of stacking to a minimum. We show some additional examples below to illustrate applied concept questions with varying degrees of stacking.

Forecasting: Thirty students are divided into two groups. Group 1 has 12 students and Group 2 has 18 students. For the next two months, each group of students will go door-to-door to sell fruit. Each group works for four consecutive days, then takes three days off before starting again. Group 1 starts working on this Monday and Group 2 starts two days later, on this Wednesday. We create a time series by recording the total number of baskets of fruit that are sold on each day. We would expect this series to display a seasonal component. How many observations (days) would we expect in each seasonal cycle?

  • a) 2

  • b) 3

  • c) 4

  • d) 7

  • e) 30

LP Sensitivity Analysis: A company uses a linear program that has the objective of maximizing the number of customers it serves this week and containing a constraint that the company may not spend more than $15,000 in providing these services. The shadow price of this constraint would be measured in

  • a) customers/dollar

  • b) dollars/customer

  • c) dollars/server

  • d) servers/customer

Both these examples provide the method-application pairing and avoid calculation. Reasoning-reading is needed (as it is in all word problems) along with memorization of the definitions of seasonal component and shadow price.

Queuing Analysis: (After a scenario involving the fitting rooms in a store.) In this scenario, which of the following would be represented by the queuing symbol Lq?

  • a) The average length of time that a customer spends either in line or using a changing room.

  • b) The average length of time that a customer waits for a changing room to become available.

  • c) The average number of customers using a changing room.

  • d) The average number of customers waiting for a changing room to become available.

  • e) The probability that a customer has to wait for a changing room to become available.

Queuing Analysis: A certain queuing system has Pw = 0.1. In this queuing system, how would W and Wq compare?

  • a) W is sure to be greater than Wq.

  • b) W is sure to be at least as great as Wq, but they might be equal.

  • c) W and Wq are sure to be equal.

  • d) Wq is sure to be at least as great as W, but they might be equal.

  • e) Wq is sure to be greater than W.

Both these examples of applied concept questions from queuing also involve reading/reasoning and memorization. Neither involves computation. The first one requires identification to link a term, and the second one uses a compare/contrast modality.

LP Modeling: Shakira is buying cookies for a party tonight and wants to have at least five cookies for each person attending. If C is the number of cookies that she buys and P is the number of people attending, this requirement could be expressed in a linear program as

  • a) C ≥ 5P

  • b) C = 5P

  • c) 5CP

  • d) 5C = P

  • e) C = 50, P = 10

This question demonstrates the error trapping modality which we discuss in Section 3.4; many students are tempted to represent “five cookies” as 5C. A student could find the correct answer simply by plugging in numbers for validation, so the instructor must decide whether to allow such shortcuts based on the targeted LO.

LP Graphical Solution: The graph below shows all of the constraints in a certain linear program, as well as an original objective function line (OFL). The x and y axes are not constraints, and the original OFL is parallel with constraint line 2. As usual, the constraint line arrows show the feasible side of the constraint and the OFL arrow shows the direction of improving objective function value.

What can we conclude about this program?

  • a) The program definitely has alternative optima.

  • b) The program is definitely a MAX program.

  • c) The program definitely has exactly one optimal solution.

  • d) The program is definitely infeasible.

  • e) The program is definitely unbounded.

This question is more complex and involves considerable stacking because deliberation on the options requires recalling the definitions of several terms, followed by the determination of whether they apply in the current context. The only “calculation” skill, however, is identifying the feasible region. The question design allows for several variations: we get an infeasible LP in the graph as shown, an unbounded LP reversing the directions of arrows 1 and 3, alternative optima by reversing all three constraint arrows, and a unique optimal solution by reversing all four arrows. Answer b is error trapping the common misconception that maximization LPs always have OFL arrows pointing up and to the right.

3.2. Make Your Scenario Descriptions and Questions Unambiguous, Concise, and Readable

Expect that your test, like any other kind of professional document, will begin as a rough draft that will probably need considerable revision before it is ready for public scrutiny. Here are several important points to keep in mind when editing your draft, followed by some examples of their application.

  • Be aware of the “curse of knowledge” (Froyd and Lanye 2008). This expression refers to the tendency to assume that others are aware of information or relationships that seem obvious to you. It can occur with even nontechnical content. For example, “the children in your family” may seem an unambiguous phrase at first. But does it refer to you and your siblings or to your own children? Does it matter if the children are no longer minors, no longer living with you, or no longer alive?

  • Organize your data and information sensibly. Events in a scenario are usually easier to follow when they are laid out in chronological order. Do not forget units in numeric data or in the definition of variables representing measurable quantities. Numeric data are usually most clearly presented when similar pieces of information appear together. If there are two natural-seeming partitions of a set of numerical data, a table may be the best way of conveying it.

  • Avoid overlong sentences. Scenarios and many stand-alone questions will have multiple parts interacting with one another. Do not try to lay out everything at once. A period at the end of a sentence gives the student a chance to stop and digest what they have been told thus far.

  • If you refer multiple times to the same item or activity, use the same word each time. Many writers vary word choice so as not to bore the reader, but this can be unwise in word problems. Repeating a word makes it clear that the item or action being discussed has not changed. In contrast, if a problem talks about both “filling boxes” and “packing cartons”, the reader is uncertain as to whether one activity is being discussed or two.

  • Use “red herrings” (Collins 2006) sparingly. Including unneeded information in a question or scenario should be a conscious choice, not an irrelevancy left over from an earlier draft. Each additional piece of information adds to the cognitive load (Sweller 2011) of the problem and the inclusion of such red herrings can lead even good students to wonder whether they have misunderstood the question.

  • Make your presentation clear and memorable. Students may be unable to correctly answer a question, but they should always be able to understand it. A scenario or question provides information to the reader, poses a question relating to that information, and finally provides choices for that answer. In a well-written MCQ, all these parts must be easily understood and easily remembered by the reader. How?

In a scenario, show your reader the situation you are describing and use active verbs to bring it to life as it unfolds; this will make it easier for students to recall it when they are grappling with individual question items. When possible, allow students to verify their understanding, such as providing an example or stating the meaning of a few of the values in a table. Consider highlighting certain items, such as words that are meant to be taken as terminology or phrases that convey an important aspect of the scenario that might easily be overlooked. If possible, have someone else read your final draft and tell you of any difficulties they found. If this is not possible, set your draft aside for a time, then read it aloud. This helps to identify any awkward or unclear phrasing.

To demonstrate these ideas, let’s look at a first draft of a scenario, one intended to assess the three probability objectives presented in Section 3.1.

First draft: There are a lot of comments about restaurants on social media and Yelp and not all of them are good, leading owners in Vienna, near Washington, District of Columbia, to see if service expectations are being met at two restaurants at the mall and near the metro station. 25% of the time, meals were eaten in the first restaurant. Ten percent of the Restaurant A customers were unsatisfied, 20% of the reviews for Restaurant B were unsatisfactory, but 50% were satisfied at both restaurants. A server is selected from one of the restaurants to see how they’re doing.

Admittedly, this initial draft may be worse than what one normally sees, but we have encountered cases where the text is only slightly less egregious. Let’s see what’s wrong with it and how we might improve the attempt.

The curse of knowledge is everywhere here. The “owners” are presumably the owners of the restaurants, but are both restaurants “at the mall and near the metro station,” or is there one at each place? Is it possible for a person to eat at both restaurants, as suggested by “50% were satisfied at both restaurants”? Are we to believe that 25% of all the meals served were served at the first restaurant, or that only 25% of that restaurant’s customers ate their meals? What are we to infer about people who were neither satisfied nor unsatisfied?

The organization of the text makes it difficult to keep track of what information is being conveyed; it goes from Restaurant A to Restaurant B to both in a haphazard and ambiguous way. Inconsistent terminology causes trouble, too. “Restaurant A” and “the first restaurant” were both used, leaving it unclear whether these are the same restaurant or different ones. Similarly, we are left to guess if there is a meaningful distinction between “unsatisfied customers” and “unsatisfactory reviews.” The first sentence is overly long and wandering, in part because it contains vague information that is irrelevant to the probability questions that we’ll be asking. All this haze can serve to obscure the fact that almost no sensible questions can be asked about this scenario since the information provided is about customers or meals or reviews, but the focus of the scenario is on “how [the server] is doing.” Let us see whether we can improve this first draft.

Revised draft: Yesterday, customers finishing their meals at Poney's Restaurant were asked to rate their respective servers. All of them did so. Based on the results, Poney's assigned to each of these servers a rating of excellent, satisfactory, or unsatisfactory. Poney's found that a server’s rating was linked to which shift, early or late, the server worked.

In the early shift, 40% of the servers received an excellent rating, 50% of them received a satisfactory rating, and the remaining 10% received an unsatisfactory rating. In the late shift, the corresponding figures were 30% excellent, 50% satisfactory, 20% unsatisfactory. Twenty-five percent of the servers working yesterday were on the early shift, and the remaining 75% were on the late shift. No server worked both shifts.

To look further into issue of service quality, Poney's interviews a server randomly selected from those who worked yesterday. For convenience, we’ll call this person Server X.

Although the new text is considerably longer than the original, it’s much easier to digest. We see a brief context, a narrative of what happened, numerical information conveyed in an orderly and unambiguous way, and consistent terminology used throughout the problem. The text focuses on the evaluation of the servers and how that evaluation is linked to the shift on which they worked. Specifying that the survey was conducted yesterday, that the restaurant is named Poney’s and that the server will be called Server X give us unambiguous ways to reference these things in the questions to come.

The numerical information could have been conveyed in tabular form if desired, especially because the data partitions naturally along both the dimensions of shift and of rating:

Table

Table

Server rating
ExcellentSatisfactoryUnsatisfactory
ShiftEarly (25% of servers)40%50%10%
Late (75% of servers)30%50%25%

If presented in this way, the text should include a statement such as “For example, 40% of the servers on the early shift received an excellent rating.” Without such an example, the scenario requires the student to determine if the probabilities are conditional on the row, conditional on the column, or joint probabilities. Although the revised scenario reflects the same math as the original, the context has shifted. It is no longer about customer reactions in two restaurants, it is now about server ratings on two shifts. Why the change from multiple restaurants to multiple shifts? We’ll discuss this in our next principle.

3.3. Find the Right Context for Your Question or Scenario

Choosing an appropriate context will increase the memorability and relevance of your question or scenario, as well as facilitate application of Principle 2 to ensure readability and clarity. Here are three specific suggestions.

  • Choose a context that is familiar to your students and is easily visualized by them. An unfamiliar context (or even unknown vocabulary) adds to the students’ cognitive load (Sweller 2011) and increases the difficulty of problems without improving their assessment value. A golfer might be tempted to write a problem with golf as its context, but why should any particular student be expected to know (or to learn during the test!) that in golf a low score is desirable, that it’s better to be “below par” than “above par,” and that the term “hole” may or may not refer to an actual hole? A familiar context often has the additional advantage of giving students the opportunity to do a “reality check” on their answers, a habit we should encourage.

  • Choose a context whose assumptions strongly echo the model requirements. Any mathematical model or formula has certain requirements for its application. Either explicitly or implicitly, the information you provide to students must allow them to reasonably assume that these requirements are met. Choosing a context in which these requirements are natural or sensible makes the problem more tractable. In contrast, contexts which implicitly contradict the model requirements will create unnecessary confusion and invite debate when the test is returned.

  • Generally, choose a context that is relevant to the material being studied. We gave the Poney’s scenario a business context for a business analytics test. The restaurant and its servers provided a setting that was easily visualized and had little chance of a curse of knowledge problem. To align the context assumptions and model requirements, we recast the original draft with two different shifts rather than two different restaurants, and we made explicit that the survey was a one-day affair. Why? Because the probability questions that we want to ask will require partitioning the population into disjoint sets in two different ways. The recasting avoided having to eliminate the possibilities that a server worked different shifts (or received different ratings) on different days, or that a customer visited both restaurants, perhaps even on the same day.

Our discussion thus far applies to both scenarios and questions, but we have primarily focused on the challenge of creating good scenario text. Now we look more closely at the task of creating the MCQs themselves.

3.4. Consider Your Distractors as Carefully as You Consider the Correct Answer

When we write the stem of an MCQ, we know which answer we intend. Surprisingly often, the other answers are written as afterthoughts and receive little of our critical attention. However, bad distractors can cause real problems (Gierl et al. 2017), so we have the following advice.

  • Avoid absurd distractors. At the very least, when a distractor completes a stem, it should result in a grammatically correct sentence. There is no reason to include a distractor that you are confident none of your students would select.

  • Verify that none of your distractors could be considered a correct answer, even an “inferior” correct answer. Unintentionally creating a second right answer to a question is remarkably easy to do, as this example shows:

    A coin is flipped two times. Let A be the event that both flips come up “heads”. Let B be the event that at least one of the two flips comes up “tails”. Then the events A and B are _______ events.

    • a) complements

    • b) comprehensively exhaustive

    • c) independent

    • d) mutually exclusive

    Answer a) cannot be correct because “are complements events” is grammatically incorrect, although students may think “complementary” was intended, which would be correct. However, the question already has two correct answers: both answers b) and d) are correct. If you’re asking students to choose the best answer to a question, you’re tacitly admitting that there are multiple correct answers. Students should not have to reproduce the reasoning or intuition that led you to decide that this correct answer is better than that one.

  • Especially in numerical problems, use distractors that provide insight into what mistakes students are making. We call this “error trapping.” An experienced teacher will know the errors students commonly make. Distractors can be created that echo these mistakes. This allows teachers to assess how often these errors are made by their students.

  • List your answers in a sensible order. List numeric answers in either increasing or decreasing order. Do the same with other ordinal quantities such as “small,” “medium,” and “large.” This helps students to find the answer they want more quickly. Put answers that differ only slightly from one another together to highlight the small but crucial differences among them. When no order seems to be preferred, alphabetical ordering can be used; it won’t aid the student, but it may keep you from putting the correct answer in your “favorite” slot. Obvious exceptions to this are true/false and yes/no questions.

  • Think twice before including answers that refer to other answers. Using “none of the above” as the last choice on a question that involves no calculation is generally okay, but avoid using it on calculation questions (Haladyna et al. 2002). Questions with choices such as “a), b) and c)” or “a) and b) only” can usually be rephrased in a way that is less ambiguous and clumsy (see Section 3.6).

To demonstrate these ideas, consider a set of MCQs for the Poney’s Restaurant scenario presented earlier, where each question addresses one of the three LOs:

  1. Consider this event: Server X receives a rating of satisfactory. The complement of this event is

    • a) Server X receives more than one rating.

    • b) Server X receives a rating of unsatisfactory.

    • c) Server X receives a rating of unsatisfactory or excellent.

    • d) Server X works the early shift.

    • e) Server X works the early shift or the late shift.

  2. We are told that 40% of the servers on the early shift received an excellent rating. Symbolically, we could express this as

    • a) P(early AND excellent) = 0.4

    • b) P(early OR excellent) = 0.4

    • c) P(early) = 0.4 and P(excellent) = 0.4

    • d) P(early | excellent) = 0.4

    • e) P(excellent | early) = 0.4

  3. Recall that Server X was selected randomly from among yesterday’s servers. What is the probability that Server X worked the late shift yesterday and received an unsatisfactory rating? Round your answer to three decimal places.

    • a) 0.075

    • b) 0.150

    • c) 0.200

    • d) 0.267

    • e) 0.900

The correct calculation for question 3 is P(late AND unsatisfactory) = P(late) × P(unsatisfactory | late) = 0.75 × 0.2 = 0.15. The distractors are P(unsatisfactory | late) = 0.20, 0.2/0.75 = 0.267 (echoing a conditional probability calculation), 0.75 × 0.1 = 0.075 (from using P(unsatisfactory | early)) and 0.75 + 0.15 = 0.90, echoing the addition law for mutually exclusive events. Doing a reality check would prevent any student from choosing this last distractor, of course, since Server X has only a 75% chance of being a late shift worker to begin with. In general, the care taken with distractors even extends to the interaction of a distractor with another distractor or the text of a different question, which we explore in the next section.

3.5. Avoid Spoilers and Double Jeopardy

Spoilers and double jeopardy arise from the interaction of available answers in one or more questions. A spoiler is a distractor that can be eliminated by logic alone, without knowledge of the course material. Either it contradicts the explanatory text in a different question or its being correct would imply that another choice would have to be correct as well. Double jeopardy occurs when correctly answering one question is dependent upon correctly answering an earlier question. In consequence, an error on the earlier question item results in losing points on both questions. In the following example from LP, we have omitted the scenario description itself.

  1. The objective function for this linear program is

    • a) x + y

    • b) 3x + y

    • c) 3x + 2y

    • d) x/3 + y

    • e) x/3 + y/2

  2. The value of the objective function when x = y = 1 is

    • a) 0.833

    • b) 1.333

    • c) 2

    • d) 4

    • e) 5

  3. How many of the program’s five constraints are redundant?

    • a) zero

    • b) one or two

    • c) two or three

    • d) three or four

    • e) four or five

  4. One redundant constraint in this program is 10x + 10y 100… (partial question only)

Question 2 is intended to test whether the student knows how to evaluate a linear expression at a point, but an incorrect answer to Question 1 will lead to the evaluation of the wrong function, and hence an incorrect answer to Question 2. One error thus results in incorrect answers to two MCQs, that is, double jeopardy. Question 3 contains spoilers: if the number of redundant constraints was two, three, or four, there would be more than one correct answer to the question. Thus, we can therefore logically eliminate answers c) and d) from consideration. Spoilers also occur when one question’s explanatory text eliminates distractors in a different question. The beginning of Question 4 tells the student that the program has at least one redundant constraint, so answer a) can be eliminated from Question 3.

3.6. Assess and Adjust the Degree of Stacking in Your Questions, and Avoid Pointless Stacking

Much of our discussion thus far has essentially focused on avoiding unintentional and unwanted barriers to a student’s comprehension of a scenario, stem, or answer choice. If we have done so successfully, the student can focus on their true task: recalling and appropriately applying the concepts, skills, and terminology of the course to move from our question to their answer, a process that will include their memory, reasoning, and judgment. In the most straightforward MCQ types, such as memorization or identification, this may be a one-step process. More complex and subtle questions, though, involve additional steps. These steps may be one or more lines of calculation but often will involve bringing a relevant fact or concept onto the “mental desktop”, then combining two or more items on that desktop in a meaningful way. We use such intentional stacking in the creation of our method-application pairing and applied concept MCQs. However, it is essential to note that there are four important consequences of higher degrees of stacking in a MC environment, even once violations of Principles 2–5 have been eliminated.

  • Stacking allows testing of chains of reasoning and higher-order thinking. This, indeed, is the primary reason for having stacking in a question. Success in the question is evidence that the student could construct such a chain or apply the relevant skills correctly.

  • Stacking increases the difficulty of the problem. The greater the degree of stacking, the more links are needed in the chain from question to correct answer. The student must keep more things in mind while forging that chain, and the number of possible arrangements of the question’s components expands combinatorically. Again, this is often what we hope to assess with intentional stacking. On the other hand, a failure in any part of this process is likely to result in an incorrect answer. Although this is also true in a CR question, no partial credit can generally be awarded for such failure on an MCQ.

  • Stacking can affect the diagnostic value of the question. Most students who get a highly- stacked question correct will have the entire skill set the question was designed to test. Such a question tells us a lot about these students. In contrast, an incorrect answer to such a question can be the result of a break anywhere in the chain of work that the problem requires. This can make it difficult to determine where and to what extent the student went wrong. Conversely, we learn relatively little about a student who gets a minimally stacked question correct, but such a question gives us clear evidence of a weakness or deficiency in students who get it wrong.

  • Stacking increases the time required to complete the problem. All other things being equal, it takes more time to do more steps. This is amplified by the fact that the student must consider more arrangements or interactions of the question components.

The previous observations show that even intentional stacking comes with a cost, which lead us to the following additional suggestions:

  • Use only the stacking you need to assess the intended objective of the question. As an example, consider a question with the prompt, “Which of the following statements is not true?” The question is perfectly clear and concise, but the stacking of such a question includes evaluating the truth value of each of the possible answers and then choosing as the right answer to the question the one statement that is wrong. Students will engage with the individual statements and will often have often completely forgotten about the “not” by the time they have finished doing so (Rodriguez and Albano 2017). If such a question is asked, it’s essential to emphasize the word “not.”

  • Consider whether a highly stacked question would be better presented as multiple, less stacked questions. By identifying the stacking inherent in the question, you can often find natural “break points” in the needed work that will allow such a decomposition. The result can be a set of MCQs that explore the depth and subtlety possible in a CR question without being overwhelming. The replacement collection effectively allows you to offer partial credit on a test with MCQs.

  • Avoid questions in which the same skill is invoked many times. A student who can compute a 3-period simple moving average forecast for one time period can (barring careless arithmetic mistakes) compute a 14-period simple moving average for five different time periods. What is the point of including the extra work on a test? As a corollary to this, do not give multiple questions that assess exactly the same skill or item.

  • Consider alternative formats. Just because a question can be asked in a MC format does not necessarily mean that it should be. Matching questions and true/false questions are two examples of alternatives that are as easily graded as the standard MCQ. Multiple true/false questions would have been a viable alternative to the “which of the following statements is not true” example given above. Additionally, most learning management systems (such as Blackboard or Canvas) provide the capability of formula questions: questions where different students are given different values for the constants in the question and must input the numerical value of the answer for that particular set of constants. This allows for the creation of an online question that naturally discourages cheating among students.

4. Examples of Applying the MCQ Design Principles

We begin with an example of a poorly written MCQ from LP.

  1. A linear program has an optimal solution of P = 60, x1 = 4, x2 = 5. One of the constraints in this program is 3x + 2y ≥ 15. In light of this information, which of the following statements is true?

    1. (4, 5) is a feasible solution to the linear program.

    2. 3x + 2y ≥ 15 is binding.

    3. 3x + 2y ≥ 15 is redundant.

      • a) I

      • b) II

      • c) III

      • d) I and II

      • e) all of the above

  2. An optimal solution to a linear program is always feasible. Is a feasible solution always optimal?

    • a) No

    • b) Yes

    • c) Sometimes

What is wrong with these questions? First, the curse of knowledge. We are supposed to infer that there are only two variables in LP and that notation was changed to make x1 into x and x2 into y. Presumably, P is the objective function value, although it is a red herring irrelevant to the problem. We suppose (4, 5) is meant to mean “x = 4, y = 5”. The author took “binding” to mean “binding on the optimal point given,” but does not say so, rendering the choice ambiguous. “Which of the following are true” may tempt some students to include statements that may be true, leading to arguments when the test is returned.

The answers to Question 1 are ambiguous, and the situation is only made worse by the grammatical error of saying “is true” rather than “are true,” implying that exactly one of the three statements is true. With this error corrected, we still have to guess that the answers a) to c) are meant to be read as “I only,” “II only,” and “III only.” However, then the answer “all of the above” is problematic. If this means “a) through d) are all true” as one would usually imagine, the answer is a spoiler, because the newly interpreted a) and b) cannot both be right. This question would be easier to digest and answer if it were broken up into three true/false questions. Question 2 has its own flaws. It spoils Question 1 by assuring that statement I is true. Putting “No” before “Yes” feels unnatural. Answer c) on Question 2 is absurd. How can something be always true “sometimes”? If we clean up the question, it might look like this:

A particular linear program has only two decision variables, x and y. The program has a unique optimal solution at x = 4, y = 5. One of the constraints in this program is 3x + 2y ≥ 15. Mark each of the following statements as true if it is certain to be true given the information provided, otherwise mark it as false.

  1. x = 4, y = 5 is a feasible solution to the linear program.

    • a) True

    • b) False

  2. 3x + 2y ≥ 15 is nonbinding on the optimal solution.

    • a) True

    • b) False

  3. 3x + 2y ≥ 15 is a redundant constraint.

    • a) True

    • b) False

What are we assessing here, and how much stacking is involved? Question 1 is one step, testing whether the student knows that an optimal solution must be feasible. Question 2 involves moderate stacking: the student needs to know that nonbinding constraints have nonzero slack, know how to compute the slack in a constraint given the coordinates of a point, and correctly perform a computation. Question 3 requires remembering what a redundant constraint is and then coming to realize that one cannot tell if a constraint is redundant without knowing the other constraints in the program. Having all of this going on in the original question format would have resulted in unnecessarily heavy cognitive load, even if the text had been appropriately edited.

We move on to examples from queuing theory. Our LOs are

Students will be able to

  • (1) Associate commonly used queuing statistics with their symbols and their meanings in a word problem.

  • (2) Identify the defining characteristics of a queuing model: population size, maximum queue length, pattern and rate of arrivals, queuing discipline, pattern and rate of service, and number of servers.

  • (3) Use queuing equations from a model to compute relevant queuing statistics: L, Lq, W, Wq, ρ, Pw,, and Pi.

These objectives are both specific and measurable. Here, we have decided to evaluate the first two objectives with a blocked-customers-cleared (BBC) model, one in which customers who are not served immediately leave the queue.

Fun-Time Kayaks is one of several kayak (boat) rental businesses that operate from a dock on a large lake. Fun-Time has 10 one-person kayaks to hire out. On average a customer keeps a kayak for 0.8 hours before returning it to Fun-Time’s dock. Those interested in renting a Fun-Time kayak arrive at its dock in a Poisson fashion with an average of five people per hour. Because of the competition from other rental agencies on the dock, however, people will not wait for one of Fun-Time’s kayaks if none is available at the time they arrive. Assume that we measure time in hours when solving this problem.

  • Unambiguous, concise, readable: The scenario avoids the curse of knowledge by not assuming that the reader knows what a kayak is, nor what time units will be used in our answers. It uses the work “dock” consistently (as opposed, for example, to sometimes calling it a “pier”). Sentences are fairly short and written in a practical style: we see a dock on a lake and people climbing into and out of one-person boats while other people are paddling about. Given the questions we intend to ask, below, there are no red herrings.

  • Context: It’s obviously a business scenario. The calling units in the problem are people, so we used boats that hold only a single person. The introduction of nearby competitors that offer the same service aligns with the model requirement that potential customers will not wait for service. We did not want students to assume exponentially distributed service times so we chose a scenario where an exponential distribution would be unlikely. The context is not perfect; for example, it is likely that groups of people will sometimes be arriving together, contrary to the Poisson arrival assumption. Still, the match is quite good. A key observation, that this system has no queue, is highlighted.

Now let us look at some possible questions that address our first two objectives. They increase in stacking as one moves down the list. One might choose only to include a subset of them on an exam.

  1. In this problem, the value of λ would be measured in

    • a) hours

    • b) kayaks

    • c) people

    • d) people per hour

  2. In modeling this queuing system, it is reasonable to assume

    • a) a finite population

    • b) an infinite maximum allowable queue length

    • c) constant interarrival times

    • d) exponentially distributed service times

    • e) ten servers

  3. Since we are solving the problem using hours as the unit of time, the numeric value of μ is

    • a) 0.2

    • b) 0.8

    • c) 1.25

    • d) 5

    • e) 12.5

  4. The scenario does not give you the equation for W in this model (in fact, it’s not a formula you’ve seen before), but you can find the value of W simply by thinking about what W means. What is this value?

    • a) 0 hours

    • b) 0.2 hours

    • c) 0.8 hours

    • d) 8 hours

  5. Some potential customers are lost because they arrive when no kayaks are available. Which of the following would tell us what fraction of potential customers are lost in this way?

    • a) P0

    • b) P1

    • c) P9

    • d) P10

    • e) PW

  6. Assume that Funtime charges customers 20 cents for each minute that they keep a kayak, and the company is interested in its average hourly revenue. Which expression would give the value of this quantity in dollars?

    • a) 4(1 – P10)

    • b) 5(1 – P10)

    • c) 12(1 – P10)

    • d) 48(1 – P10)

    • e) 60(1 – P10)

    For the first question, the student needs only know that λ is the mean arrival rate of a queuing system; that is, mean arrivals per unit time. The second question deals with five different queuing system characteristics but knowing what a “server” is would be sufficient to answer the question correctly. For Question 3, the student must know that μ is the mean service rate per server, find the stated service time in the problem, and then convert that time to a rate. The distractors include 1/μ, 1/λ, λ, and nμ.

    Question 4 requires that the student know that W is the mean time a calling unit spends in the system, that this system has no queue, that this in turn implies that W is the time mean time in service, and that the problem gives mean service time to be 0.8 hours. Question 5 requires the student to realize that Pi is the probability that the queuing system contains i calling units, that a newly arriving customer is lost if and only if there are 10 customers already in the system, and that with Poisson arrivals, the probability of the system being in a given state is the same as the probability of the system being in that state when a new arrival appears. In addition, the student should know that PW has no direct connection to W itself. That is a higher degree of stacking, perhaps, than was originally evident, although knowing the meaning of P0 and P1 should be enough to eliminate them from consideration.

    Question 6 is offered as an example of a question that is almost certainly too hard for an introductory analytics course. The stacking includes knowing what P10 represents in general, what it means in this problem and therefore what 1 – P10 means here, knowing that when arrivals are Poisson the probability of a full system is the same as the probability of a lost potential customer, and finally the calculation of how much money the company would make per hour if no potential customers were lost. If the instructor wanted to evaluate the ability of students to answer Question 6 at an introductory level, it would be more appropriate to break it up into subquestions: “How much on average does a customer who gets a kayak pay?,” “What would the expected revenue be if every potential customer got a kayak?,” and so forth.

    The kayak scenario is not good for testing the third objective, because the equations for a BCC queuing system are mathematically complex. However, a stand-alone question could be used to assess it:

  7. JazzyWash is an automatic car wash capable of cleaning one car at a time. Every car takes exactly 5 minutes to clean, so the car wash is capable of handling up to 12 cars per hour. Customers arrive there according to a Poisson distribution with a mean arrival rate of 5 per hour and wait in line until it is their turn for their car to be washed. In such a queuing system the equation for Lq is λ2σ2+(λμ)22(1λμ). Compute the value of Lq for JazzyWash.

    • a) −2.057

    • b) 0

    • c) 0.149

    • d) 2.005

    • e) 21.578

Distractors here include reversing μ and λ (or making them times rather than rates), using σ = 1, the most common line length, and a credible looking value, 2.005. Any student choosing either of the first two answers does not understand the meaning of Lq. The difficulty of this question could be scaled up by giving mean interarrival times instead of λ or not supplying μ. It could be made easier by explicitly stating that λ = 6 cars/hr, μ = 12 cars/hr, and σ = 0 hr.

We round out this section with one more example of a good scenario, this one in decision analysis and involving LOs for constructing a tree, rolling back a tree, computing probabilities for research branches, and the concepts of expected value of sample information (EVSI) and expected value of perfect information (EVPI). To provide more real-world context, one could discuss why UPS drivers almost never turn left on their delivery routes. The reader is invited to examine the scenario from the perspective of the six principles and judge how well it adheres to them. Although it is a lengthy scenario (a full page), it is also rich and could be used to test most if not all LOs in decision theory at the introductory level.

The map below shows a system of roads connecting an apartment complex with a school. Five locations (A, B, C, D and E) are also shown. Straight Street (in blue) runs in a straight line across the top of the map connecting the apartment complex (A) to the school (E). It consists of four 1-block long sections labeled Sections AB, BC, CD and DE. The map shows two additional streets. Short Street is a four-block-long direct connection between locations B and D. Long Street is an 8-block-long direct connection between points A and E.

Jan, a college student, uses this road network to drive from the apartment complex to the school. Unfortunately, recent roadwork on Sections CD and DE of Straight Street sometimes closes one or both of those two sections. Jan will not know if one of these road sections is closed until arriving at that section. If her route is blocked by a closed section, she’ll have to backtrack and take a different route to school. Road section CD has a 60% chance of being closed. Road section DE has a 10% chance of being closed. Whether one of these sections is closed is independent of whether the other is closed.

Example: Jan drives along Sections AB and BC of Straight Street, only then discovering that Section CD is closed. She must therefore turn around and return to point B to choose a different route. Her drive up to the point when she returns to B will be three blocks long, and she’ll still need to make her way to E from there.

The rolled back decision tree for determining the strategy that allows Jan to minimize her average travel distance (in blocks) is shown below. Note that three of the values in the tree (labeled as Q1, Q2 and Q3) are missing. You’ll supply them in the first three questions.

  1. The value at location Q1 on the tree should be

    • a) 4

    • b) 7.12

    • c) 7.32

    • d) 7.52

    • e) 20

  2. The value at location Q2 on the tree should be

    • a) 7.12

    • b) 16

    • c) 17

    • d) 17.2

    • e) 18

  3. The value at location Q3 on the tree should be

    • a) 4

    • b) 6

    • c) 9.2

    • d) 10

    • e) 12

  4. If Jan follows the optimal strategy, it is possible that she may drive exactly ___ blocks.

    • a) 4

    • b) 5

    • c) 6

    • d) 8

    • e) 20

  5. Suppose Jan faces this same situation over many successive days. The value of 7.12 over the leftmost node in the tree implies that if Jan follows her optimal strategy, the length of her trip to school

    • a) will always be exactly 7.12 blocks.

    • b) will always be either 7 blocks or 8 blocks.

    • c) will never be less than 7.12 blocks.

    • d) will never be more than 7.12 blocks.

    • e) will, on average, be 7.12 blocks.

  6. For this question only, suppose that Jan began her trip at Point B and ended her trip at Point D. As before, her goal is to minimize the expected number of blocks she must drive. The simplest decision tree for the modified problem has

    • a) one chance node, followed by one decision node.

    • b) one decision node, followed by one chance node.

    • c) two chance nodes, no decision nodes.

    • d) two decision nodes, no chance nodes.

    • e) two decision nodes, two chance nodes.

    This example demonstrates how a basic structure may be adapted to test different LOs at differing levels of complexity. Questions 1 and 2 involve minimal stacking and test the mechanical evaluation of a decision node and a chance node. Question 3 requires the student to understand the sequence of events represented by the path leading to the desired payoff and then to use the map to compute the number of blocks on Jan’s route. This involves some stacking and is a more sophisticated question. Question 4 stacks identifying the optimal strategy with the meaning of chance nodes, decision nodes, and payoffs. Question 5 tests whether the student understands the meaning of an expected value at a node. Question 6 assesses the ability of the student to analyze a simplified version of the scenario on their own and to build the corresponding tree.

    Additional questions could be added to expand the range of LOs covered. For instance, we could ask for the probability that both sections CD and DE are open (joint probability of independent events). Alternatively, by modifying the assumption that the section closures are independent, we could test the posterior probability calculations common in research branches:

  7. For this question only, suppose we learned that one piece of information in the scenario was in error. While there is still a 60% chance that Section CD is closed and a 10% chance that Section DE is closed, the closures are no longer independent. In fact, whenever DE is closed, CD is closed as well. Given this information, if CD is closed, what is the probability that DE is closed as well? Give your answer to two decimal places.

    • a) 0.06

    • b) 0.10

    • c) 0.17

    • d) 0.50

    • e) 0.60

    This is a full-blown Bayes problem and thus involves a considerable amount of stacking.

    Returning to the original problem, the calculation of EVPI could be tested by providing the joint probabilities of the four possible open/closed combinations for Sections CD and DE and ask for EVPI for this tree. To test the understanding of EVPI without such calculation, we could ask for the EVPI of a tree representing the return trip from school to the apartment complex. (The EVPI is zero, because Jan has all of the information she needs about a road closure by the time that any decision needs to be made.) To test understanding of EVSI, we could ask the following MCQ:

  8. Suppose Jan could learn at the beginning of her trip whether Section DE were open or closed. The EVSI for this information would be 0.92 blocks. This means that if she had this information, her average optimal trip length would be

    • a) increased by 0.92 blocks.

    • b) increased by 0.92 blocks, but only if Section DE were closed.

    • c) exactly 0.92 blocks.

    • d) reduced by 0.92 blocks.

    • e) reduced by 0.92 blocks, but only if Section DE were open.

In summary, scenarios like this one provide a rich environment for asking nontrivial MCQs at different levels of difficulty.

We close this section by presenting two examples of MCQs from our classroom experience that were problematic and discussing how we used our design principles to address their respective issues. The first one was written for assessment testing in our introductory business statistics course:

Original version: The five members of a company’s board vote for either Carlos or Jasmin as new chairperson, and a simple majority wins. Each member votes, and each member has a nonzero chance of voting for either candidate. Consider these events:

  1. Jasmine wins.

  2. Jasmine loses.

  3. Jasmine gets two or more votes.

  4. Jasmine gets two or less votes.

Which of the following statements is NOT correct?

  • a) I and II are complementary events

  • b) I and III are dependent events.

  • c) I and IV are mutually exclusive events.

  • d) III and IV are collectively exhaustive events.

  • e) II and IV are independent events.

This original version already benefited from applying some of the design principles. Explicitly stating that a simple majority is required for victory avoids a curse of knowledge issue. The additional information provided prevents more than one right answer. The requirement that each member votes for one of the two candidates prevents ties from abstention or write-in votes. If at least three board members had a zero probability of voting for a particular candidate, then I and III both have probability 1 or both have probability 0, and so would be independent. Despite this, the question has serious issues. The large degree of stacking made it a difficult problem and, as a consequence, also made it of limited value in discerning what probability concept the students actually understood. Performance improved by 20% when we revised it to require comparison of only two events:

Revised version: The five members of a company’s board vote for either Carlos or Jasmin as new chairperson, and a simple majority wins. Each member votes, and each member has a nonzero chance of voting for either candidate.

Consider these two events: Jasmin wins and Jasmin loses. These two events are NOT

  • a) complementary

  • b) collectively exhaustive

  • c) independent

  • d) mutually exclusive

We have tried a variety of MCQs to evaluate LP modeling skills, including this one:

FAO, the Food and Agriculture Organization of the UN, has a list of kinds of food that it could supply to victims of a drought. The list records the cost per pound of each kind of food and the amount of each essential nutrient that a pound of that food contains. Using foods on this list, FAO is going to design a food package that can be assembled and delivered to the drought victims. One package must provide a victim with all of the essential nutrients he or she needs for one day. FAO wishes to make and deliver as many such packages as it possibly can, but it has a limited amount of money available for this project. FAO models their problem as a linear program. Which of the following choices best summarizes what the structure of this program is likely to be?

  • a) The objective is to maximize deliveries while minimizing cost. Decision variables for how much of each nutrient is in a pound of each food. Constraints about nutrients.

  • b) The objective is measured in dollars. One variable for each kind of food. One constraint for each type of nutrient, and one constraint about deliveries.

  • c) The objective measured in dollars. One variable for each kind of nutrient. One constraint for each type of food, and one constraint about deliveries.

  • d) The objective measured in deliveries. One variable for each kind of food. One constraint for each type of nutrient, and one constraint about money.

  • e) The objective measured in deliveries. One variable for each kind of nutrient. One constraint for each type of food, and one constraint about money.

This problem has an extensive text and requires students to pick out the objective, constraints, and decision variables for the LP model without developing the mathematical formulation. It had an error rate of 41%, with answers A, B, and E all being common wrong choices. We recognized it as an example of MCQ that could benefit from unstacking and replaced it with the four MCQs in the Leisure Time Scenario (see Appendix). Additional examples of MC items written for statistics and simulation analysis using our design principles are provided in the Appendix.

5. Concluding Remarks

The primary purposes of testing in the academic setting are generally viewed as the assessment of student knowledge and skills and the assignment of course grades. Yet tests also offer a host of positive benefits with respect to improving student learning by requiring practice retrieving information from memory and transferring knowledge to new contexts. For instructors, test results serve to document gaps in student understanding of complex material and provide feedback to inform teaching (Roediger et al. 2011). Test questions in MC format can be particularly helpful for identifying specific misunderstandings of analytics concepts and methods and for tracking student progress over time on accreditation standards (Wind et al. 2019, Palocsay et al. 2020). We were motivated to experiment with MCQs for all these reasons, as well as practical matters related to larger class sizes and online and hybrid class delivery modes during the pandemic. These experiments led us to articulate the framework presented herein, consisting of a set of evidence-based design principles to guide construction of high-quality MCQs for our introductory business analytics courses.

Our focus here has been on writing individual MC items that evaluate student learning in analytics across a spectrum of cognitive levels, with an emphasis on quantitative reasoning and analysis skills. However, having written good MCQs, we note that there are still a number of other important issues for instructors to consider during the process of composing and administering an effective test with MC items, including the following:

Although exploration of these issues is beyond the scope of this paper, we hope that this initial listing will encourage discussion and stimulate future research on test writing with MCQs for analytics education. Most of the relevant work published to date on these topics has been in other disciplines such as psychology and medical education. There are many open questions about how their empirical findings may (or may not) transfer to introductory OR/MS/analytics courses that are taught to students with different mathematical backgrounds in a variety of business and engineering programs. Investigation of their impact on student achievement in analytics can lead to better instruction and improved learning outcomes.

Appendix. Additional Examples of MCQs

Questions 1–4 deal with the Leisure Time scenario, described here.

Leisure Time, Inc. makes and sells patio chairs and patio tables. Each chair costs $15 to make and sells for $35. Each table costs $50 to make and sells for $180. Because of research that Leisure Time has done into the buying habits of its customers, the company has decided that at least 80% of the items it manufactures must be patio chairs.

Leisure Time will not produce units that it cannot sell, so the number of an item made will be the same as the number of that item sold. Without advertising, Leisure Time knows that its customers generate a demand for 3,000 chairs and 800 tables. To increase the number of units demanded, Leisure Time may advertise. Each ad that the company runs costs $100 and increases the demand for Leisure Time products by one table and three chairs. Whether it advertises or not, the company is not obligated to completely satisfy the demand.

Leisure Time has a total operations budget of $228,500, which it can spend on making furniture and advertising. Its goal is to maximize its profit in dollars from its production and sales.

We represent this problem as a linear program in which T is the number of patio tables made and sold, C is the number of chairs made and sold, and A is the number of ads run.

  1. The objective function for this linear program would be

    • a) Maximize 20 C + 130 T

    • b) Maximize 20 C + 130 T – 100 A

    • c) Maximize 20 C + 130 T – 100 A – 228,500

    • d) Maximize 35 C + 180 T

    • e) Maximize 35 C + 180 T – 100 A

  2. The requirement that at least 80% of the items manufactured by Leisure Time must be chairs is represented by the constraint

    • a) 0.80 CT

    • b) C ≥ 0.80(C + T + A)

    • c) C ≥ 0.80(C + T)

    • d) C ≥ 0.80 T

    • e) C ≥ 8000

  3. The constraint that Leisure Time cannot sell more chairs than people are willing to buy is represented by the constraint

    • a) 35 C ≤ 3000

    • b) C ≤ 3000 + 3 A

    • c) C ≤ 3000 + 100 A

    • d) 15 C ≤ 3000 + A

    • e) 20 C ≤ 3000 + A

  4. One constraint in this program is that Leisure Time cannot sell more patio tables than the quantity demanded. In the optimal solution, this constraint is nonbinding. Which of the following conclusions follows from this fact?

    • a) Leisure Time made and sold exactly 800 patio tables.

    • b) Leisure Time sold no patio tables at all.

    • c) Leisure Time didn’t make enough patio tables to completely satisfy the demand for patio tables.

    • d) Leisure time made exactly enough patio tables to exactly satisfy the demand for patio tables.

    • e) Leisure Time made more patio tables than people were willing to buy.

End of the Leisure Time scenario.

Questions 5–7 deal with the Auction scenario:

You are offering a signed photograph of a famous person in a sealed-bid auction. In such an auction, bids are submitted secretly, so no bidder knows how much anyone else bid until the auction is over. You have told the potential bidders that you will not sell the photo for less than $30 and that you are only accepting bids that are multiples of $5. High bid (of at least $30) wins, and if there is a tie for the highest bid, you will randomly choose who wins from among the highest bidders. The number of people who bid on the photograph and the amount that any given bidder will bid are random variables whose distributions are given in the table.

Table

Table

No. of biddersProbabilityAmount bidProbability
00.1$300.31
10.1$350.36
20.5$400.29
30.3$450.04

If no one bids at least $30 for the photo at the auction, you will sell it to a friend for an agreed upon price of $20. We are interested in the expected (average) amount of money that you’ll make from the sale of the photo.

  • 5. Which of the following statements about the simulation of this situation is correct?

    • a) A single run of this simulation would allow us to answer our question of interest.

    • b) A single run of this simulation might use the # of bidders process generator multiple times.

    • c) A single run of this simulation might not use the amount bid process generator at all.

    • d) The expected amount of money you make from the sale of the photo will be one of these values: $0, $30, $35, $40, or $45.

    • e) The simulation would require a third process generator that determines who gets the picture in the event of a tie for the highest bid.

  • 6. Suppose that, following the procedure developed in class, we use the random number 0.3700 to determine the size of a particular bid. The bid corresponding to this random number would be

    • a) $30

    • b) $35

    • c) $40

    • d) $45

  • 7. I conducted an Excel simulation of this auction, and to compute how much money you make from the photo, I used this statement: =IF(MAX(E9:G9) ≥ 30, MAX(E9:G9), $B$7). Cells E9 through G9 held the size of the bids made, if any. What did cell B7 contain in my simulation?

    • a) The minimum bid that would be accepted, $30.

    • b) The number of bidders.

    • c) The highest bid made.

    • d) The amount of money your friend offered for the picture, $20.

    End of Auction scenario

  • 8. We gather data on how often Dr. Smith’s students came to her office hours during the last semester. 84% of these students came 4 times or less, 94% of them came 5 times or less, and a 98% of them came 6 times or less. Given this information, what percentage of Dr. Smith’s students came to her office hours exactly 6 times during last semester?

    • a) 2%

    • b) 4%

    • c) 6%

    • d) 10%

    • e) 16%

  • 9. The graph below shows the distributions of two normally distributed random variables, X and Y. The heavy line shows the distribution of X. The lighter line shows the distribution of Y. How do the means and standard deviations of these two variables compare?

    • a) X has the greater mean and the greater standard deviation.

    • b) Y has the greater mean and the greater standard deviation.

    • c) X has the greater mean, but Y has the greater standard deviation.

    • d) Y has the greater mean, but X has the greater standard deviation.

    • e) X has the greater mean and both X and Y have the same standard deviation.

  • 10. (For a quiz in which no tables or calculators are allowed.) 80% of all observations for a normally distributed variable have z-scores less than z = 0.84. Find P(0 ≤ z ≤ 0.84). (You may wish to draw a picture.)

    • a) 0.1

    • b) 0.2

    • c) 0.3

    • d) 0.4

    • e) 0.8

References

  • Anderson LW, Krathwohl DR (2001) A Taxonomy for Learning, Teaching, and Assessing: A Revision of Bloom’s Taxonomy of Educational Objectives (Longman, New York).Google Scholar
  • Bloom BS (1956) Taxonomy of Educational Objectives: Handbook 1: Cognitive Domain (McKay, New York).Google Scholar
  • Burns ER (2010) “Anatomizing” reversed: Use of examination questions that foster use of higher order learning skills by students. Anatomy Sci. Edu. 3(6):330–334.CrossrefGoogle Scholar
  • Butler AC (2018) Multiple-choice testing in education: Are the best practices for assessment also good for learning? J. Appl. Res. Memory Cognition 7(3):323–331.CrossrefGoogle Scholar
  • Butler-Henderson K, Crawford J (2020) A systematic review of online examinations: A pedagogical innovation for scalable authentication and integrity. Comput. Edu. 159:104024.CrossrefGoogle Scholar
  • Collignon SE, Chacko J, Wydick MM (2020) An alternative multiple‐choice question format to guide feedback using student self‐assessment of knowledge. Decision Sci. J. Innovative Edu. 18(3):456–480.CrossrefGoogle Scholar
  • Collins J (2006) Writing multiple-choice questions for continuing medical education activities and self-assessment modules. Radiographics 26(2):543–551.CrossrefGoogle Scholar
  • Downing SM (2005) The effects of violating standard item writing principles on tests and students: The consequences of using flawed test items on achievement examinations in medical education. Adv. Health Sci. Edu. Theory Practice 10(2):133–143.CrossrefGoogle Scholar
  • Froyd J, Layne J (2008) Faculty development strategies for overcoming the “Curse of Knowledge.” Proc. 38th Annual Frontiers in Edu. Conf. (IEEE, Piscataway, NJ).Google Scholar
  • Ghidinelli M, Cunningham M, Monotti IC, Hindocha N, Rickli A, McVicar I, Glyde M (2021) Experiences from two ways of integrating pre-and post-course multiple-choice assessment questions in educational events for surgeons. J. Eur. CME 10(1):1–11.CrossrefGoogle Scholar
  • Gierl MJ, Bulut O, Guo Q, Zhang X (2017) Developing, analyzing, and using distractors for multiple-choice tests in education: A comprehensive review. Rev. Edu. Res. 87(6):1082–1116.CrossrefGoogle Scholar
  • Gronlund NE (2004) Writing Instructional Objectives for Teaching and Assessment, 7th ed. (Pearson/Merrill/Prentice Hall, Upper Saddle River, NJ).Google Scholar
  • Haladyna TM, Downing SM (1989) A taxonomy of multiple-choice item-writing rules. Appl. Measures Edu. 2(1):37–50.CrossrefGoogle Scholar
  • Haladyna TM, Rodriguez MC (2013) Developing and Validating Test Items (Routledge, New York).CrossrefGoogle Scholar
  • Haladyna TM, Rodriguez MC (2021) Using full-information item analysis to improve item quality. Edu. Assessment 26(3):198–211.CrossrefGoogle Scholar
  • Haladyna TM, Downing SM, Rodriguez MC (2002) A review of multiple-choice item-writing guidelines for classroom assessment. Appl. Measures Edu. 15(3):309–333.CrossrefGoogle Scholar
  • Kuechler WL, Simkin MG (2010) Why is performance on multiple‐choice tests and constructed‐response tests not more closely related? Theory and an empirical test. Decision Sci. J. Innovative Edu. 8(1):55–73.CrossrefGoogle Scholar
  • Lawlor KB, Hornyak MJ (2012) Smart goals: How the application of smart goals can contribute to achievement of student learning outcomes. Developments Bus. Simulation Experient. Learn. 39:259–267.Google Scholar
  • Lemons PP, Lemons JD (2013) Questions for assessing higher-order cognitive skills: It’s not just Bloom’s. CBE Life Sci. Edu. 12(1):47–58.CrossrefGoogle Scholar
  • Madaus GF, O’Dwyer LM (1999) A short history of performance assessment: Lessons learned. Phi Delta Kappan 80(9):688–695.Google Scholar
  • Manoharan S (2019) Cheat-resistant multiple-choice examinations using personalization. Comput. Edu. 130(March):139–151.CrossrefGoogle Scholar
  • McTighe J, Wiggins G (2012) Understanding by Design Framework(Association for Supervision and Curriculum Development, Alexandria, VA).Google Scholar
  • Meda L, Swart AJ (2018) Analysing learning outcomes in an Electrical Engineering curriculum using illustrative verbs derived from Bloom’s Taxonomy. Eur. J. Engrg. Edu. 43(3):399–412.CrossrefGoogle Scholar
  • Miltenburg J (2019) Online teaching in a large, required, undergraduate management science course. INFORMS Trans. Edu. 19(2):89–104.LinkGoogle Scholar
  • Momsen JL, Long TM, Wyse SA, Ebert-May D (2010) Just the facts? Introductory undergraduate biology courses focus on low-level cognitive skills. CBE Life Sci. Edu. 9(4):435–440.CrossrefGoogle Scholar
  • Morrison S, Free KW (2001) Writing multiple-choice test items that promote and measure critical thinking. J. Nursing Edu. 40(1):17–24.CrossrefGoogle Scholar
  • Palocsay SW, Stevens SP, Novoa LJ (2020) STRATA: A spreadsheet tool for multidimensional analysis of operations research/management science assessment test data. INFORMS Trans. Edu. 21(1):41–56.Google Scholar
  • Paniagua MA, Swygert KA (2016) Constructing Written Test Questions for the Basic and Clinical Sciences (National Board of Medical Examiners, Philadelphia).Google Scholar
  • Roediger III H, Putnam A, Smith M (2011) Ten benefits of testing and their applications to educational practice. Psych. Learn. Motivation 55:1–36.Google Scholar
  • Rodriguez MC (2003) Construct equivalence of multiple‐choice and constructed‐response items: A random effects synthesis of correlations. J. Edu. Measures 40(2):163–184.CrossrefGoogle Scholar
  • Rodriguez MC (2005) Three options are optimal for multiple‐choice items: A meta‐analysis of 80 years of research. Edu. Measures 24(2):3–13.CrossrefGoogle Scholar
  • Rodriguez MC, Albano AD (2017) The College Instructor’s Guide to Writing Test Items: Measuring Student Learning (Routledge, New York).CrossrefGoogle Scholar
  • Scully D (2017) Constructing multiple-choice items to measure higher-order thinking. Practical Assessment Res. Evaluation 22:4.Google Scholar
  • Shin J, Bulut O, Gierl MJ (2020) The effect of the most-attractive-distractor location on multiple-choice item difficulty. J. Experiment. Edu. 88(4):643–659.CrossrefGoogle Scholar
  • Simkin M, Kuechler W (2005) Multiple-choice tests and student understanding: What is the connection? Decision Sci. J. Innovative Ed. 3(1):73–98.Google Scholar
  • Swart AJ (2010) Evaluation of final examination papers in engineering: A case study using Bloom’s Taxonomy. IEEE Trans. Edu. 53(2):257–264.CrossrefGoogle Scholar
  • Sweller J (2011) Cognitive load theory. Psych. Learn. Motivation 55:37–76.CrossrefGoogle Scholar
  • Vanderbilt A, Feldman M, Wood I (2013) Assessment in undergraduate medical education: a review of course exams. Medical Edu. Online 18(1):20438.CrossrefGoogle Scholar
  • Wang L (2019) Does rearranging multiple‐choice item response options affect item and test performance? ETS Research Report Series, ETS RR-19-02, Educational Testing Service, Princeton, NJ.Google Scholar
  • Wind S, Alemdar M, Lingle J, Moore R, Asilkalkan A (2019) Exploring student understanding of the engineering design process using distractor analysis. Internat. J.STEM Ed. 6(1):1–18.Google Scholar
  • Zhai X, Li M (2021) Validating a partial-credit scoring approach for multiple-choice science items: An application of fundamental ideas in science. Internat. J. Sci. Edu. 43(10):1640–1666.CrossrefGoogle Scholar