September 3, 2026 in Operations research
Not All Time Is Created Equal
In Firefighting, Time-Based Metrics Can Miss the Real Work
SHARE: PRINT ARTICLE:
https://doi.org/10.1287/orms.2026.03.03
Not all hours feel the same. Imagine two flights with the same gate-to-gate time. One was smooth and uneventful. The other was diverted for weather, hit heavy turbulence, and kept everyone seated. Both flights had the same duration, but the experience was very different.
This example illustrates a flaw in any operational measure that accounts for workload as a fraction of time. The denominator is fixed, and each minute is evaluated equally. But in environments in which intensity fluctuates – such as in clinical or outdoor settings, situations affected by weather conditions, or circumstances involving customer complexity – the average span of time can hide information. In the U.S. fire service, this measure is called Unit Hour Utilization (UHU), the percentage of a 24-hour shift spent on incidents.
Whether the incident occurred at 3am or 2pm, that time doesn’t reflect the human cost or the incident’s complexity. The UHU metric, adapted from ambulance economics, was not designed to measure workload, although it is used that way, as no better option exists. The study described below developed a more accurate measure and tested it with real data from a large, mid-Atlantic fire department. The measure used can apply to any operation in which work intensity varies and utilization influences staffing.
Same Gap, Different Industries
Companies that don’t understand or account for the impact of different types of working conditions on employees eventually experience high turnover or error rates. The fire service has been paying for this oversight for years, resulting in serious health consequences that follow firefighters into retirement. A weighted measure won’t change those outcomes on its own, but it may help if it gives leadership the data to make better-informed decisions.
The same problem appears in other industries. For example, a call center handling routine billing at 85% is not the same operation as one handling complaint escalations at the same rate. An intensive care unit (ICU) at 80% occupancy with three critical patients is not the same as one with three post-operative patients at 80% occupancy. The dashboards don’t show the difference, but employees can feel it.
Variability in workload intensity is a common challenge in operations research. The real difficulty lies in creating a measure that correctly captures the variability from the data that operations already gather.
Beyond Unit Hour Utilization
In the initial phase of our study, we interviewed 16 fire service data experts from seven states. Most used UHU, although they thought it inaccurately depicted firefighters’ true workload. Insights from these interviews and existing research on firefighter fatigue helped us identify four operational factors that contribute to cumulative fatigue: night wakeups, burst, volume, and call severity.
In our study, we represented each factor as a multiplier applied to the call duration. The exact multiplier values are structured assumptions rather than validated physiological constants. This approach is consistent with established methods for constructing composite indicators, which permit transparent, judgment-based weighting when no objective standard exists. We made the factors intentionally conservative. We conducted sensitivity analysis to ensure the results were robust and not driven by a single arbitrary set of weights. They multiply rather than add because, in the field, conditions compound rather than simply accumulate.
The data set we used covered 133,682 incident records across 38 operational units over roughly 640 days. Every call in the data set received a combined multiplier of 1.0 when no fatigue conditions were met, increasing as fatigue conditions stacked. A single call can reach 4.92 when high-acuity, overnight, back-to-back, and high-volume conditions coincide. Baseline workload calculations also incorporated the daily equivalent of 192 annual hours of Insurance Services Office (ISO) training and approximately two hours of physical training per shift, reflecting routine operational responsibilities beyond emergency response.
Combined, these activities account for an average of 3.58 hours of baseline workload during each 24-hour shift. Rather than treating every hour of a 24-hour shift as equally available for incident workload, these baseline responsibilities were incorporated into the workload calculation to better represent the operational demands placed on firefighters. The multiplier is the product of four factor multipliers:
Adjusted Workload = Call Duration × Night Wake-up × Burst × Volume × Severity.
Workload Calculated
- Night Wake-up: Calls from 10pm to 6am trigger a coefficient that grows with each successive interruption: 1.25 for the first wake-up, 1.50 for the second, 1.75 for the third, plus 0.25 for every additional disturbance. Sleep debt is not a step function. Each interruption costs more than the last, and recovery within a single shift does not occur.
- Burst: When a unit leaves one call and is dispatched to another within 15 minutes, the second call is flagged as a “burst.” Daytime bursts carry a coefficient of 1.25, and overnight bursts, 1.50. The factor accounts for rehydration, decontamination, report writing, and rest time lost due to back-to-back assignments.
- Volume (per 24-hour shift): Volume is binned by the cumulative call count: no adjustment up to five calls, 1.05 above five, 1.25 above ten, 1.50 above 15, and 1.75 above 20 calls. Volume also serves as a proxy for the volume of administrative work (report writing, equipment checks, restocking) that scales with call count and is invisible to UHU.
- Severity: Two incident types were used: structure fires with more than 90 minutes of on-scene time and cardiac arrests with documented compressions. The severity of structure fires increases with time on scene, from 1.50 at 90 minutes to 2.50 after 3 hours. For cardiac arrests, severity is set at 1.50 for adults and 3.00 for children. These values account for both bodily stress and the mental impact of resuscitation, which is significant with pediatric patients. When none of these four conditions are met, the coefficient defaults to 1.0, making the weighted measure standard. Differences between the two occur only when the operating environment includes conditions UHU was never built to register.
The compounding effect of the work becomes clearer through an example. Consider Engine 23, dispatched to a structure fire at 2:30am for an incident lasting 140 minutes. Three factors influence the timing: it’s the twelfth call of the shift, the crew’s second wake-up, and eight minutes after the last call.
Standard UHU says the crew spent 140 minutes on the call. The weighted measure tells a different story: the call was equivalent to 689 minutes. In this example, the multiplier is 4.92, the product of severity (1.75 for a structure fire above the two-hour threshold), burst (1.50 for an overnight burst), volume (1.25 for the 10-call threshold), and night wake-up (1.50 for the second wake-up). No single factor drives the gap. The compounding does.
The Burden of 133,682 Incidents
Burden that cannot be measured cannot be managed. We built the weight to make the unmeasured visible, and across the 38 units it did. Standard UHU spanned 3.4-30.9%. Weighted utilization spanned 4.6-43.6%. The gap between the two measures averaged 38.5%, reaching 58% in the most affected unit. Every unit moved up. Some moved up far more than others.
The people doing the work already know what leadership doesn’t see. We created a multiple regression model to analyze workload contribution. The results helped explain which factors contributed to the gap between standard UHU and adjusted workload. It did not validate the model, and it should not be interpreted that way. We selected the variables because the literature and subject-matter expertise already support their connection to workload, fatigue, recovery, and operational burden. The regression simply showed which parts of the model contributed most once those factors were accounted for.
The four factors accounted for 83.4% of the variance in weighted and standard utilization. Adding the 15 possible two-way interaction terms and retaining those supported by stepwise selection and AIC/BIC criteria increased the explained variance to 95.6%. The addition was statistically significant (F=6.9, p<0.001), confirming that these workload drivers operate jointly rather than independently. The strongest single interaction was volume×cardiac arrest (p=0.04); however, because the predictors are correlated and the sample is small relative to the number of terms, we interpreted the interactions as a group rather than by assigning a direction to any individual coefficient.
Because the interaction model packed many terms into a relatively small sample, we took several validation steps to ensure the results were not an artifact of multicollinearity, which is the tendency of correlated predictors to distort one another’s apparent effects. Centering the predictors before building the interaction terms reduced the maximum variance inflation factor from more than 228,000 to 37, and leave-one-out cross-validation confirmed that the four-factor model generalized to data it had not seen, explaining 71.7% of the variance out of the sample. The added interaction terms improved out-of-sample accuracy by only 2.2 percentage points, despite a 12-point in-sample gain, so we treated them as exploratory and retained the simpler four-factor model for inference.
To identify which factors mattered most without being misled by their correlations, we used a dominance analysis (a Shapley-value decomposition that is unaffected by multicollinearity) to fairly apportion the
model’s explanatory power. High-volume shifts and nighttime wake-ups together accounted for roughly three-quarters of the measured workload gap (40.8% and 31.8% of the explained variance, respectively), followed by structure-fire exposure at 17.1%. This ordering held, regardless of how we entered the correlated predictors, giving a stable picture of the primary drivers of hidden workload.
To enable comparison across schedules, a Schedule Exposure Index (SEI) scales weighted utilization by total exposure relative to a 42-hour reference week. A work schedule does not reflect the burden of individual calls, but it can increase how often firefighters are exposed to conditions that cause fatigue. The SEI reflects the reality that the schedules themselves carry workload and should be accounted for. The 56-hour cycle multiplies by 1.33. Once exposure is included, basic life support transport units (already busy due to raw UHU) dominate the high-burden ranking. Trucks and rescues, with low call counts but high incident severity, sit at the bottom. The top-to-bottom range stretches roughly 10-fold. UHU shows the time committed. The weighted index shows how much more the busiest units carry.
Healthcare Applications
We tested the most data-intensive version of the model, but most departments don’t have the ability to do so. The model is customized to scale for capabilities at various levels. The basic tier considers call counts, incident flags,
and wake-up counts. The intermediate tier adds night wakeups, burst detection, and volume binning. The advanced tier, which we tested, includes severity weighting and the schedule index. The biometric-enhanced tier further incorporates heart rate variability, sleep tracking, stress markers, and cognitive readiness scores at the individual level.
Similar fields within healthcare may also benefit from a weighted workload measure, although the sources of fatigue will differ across professions. While the factors contributing to fatigue vary by operational setting, cumulative workload is experienced across many high-demand environments. For non-fire operations, the same tier structure applies.
For example, an ICU might begin by grouping patients by acuity and adding a night-admission factor, then incorporate handover volume and case complexity. Private ambulance services could account for interfacility transfers, long-distance transports, and critical care responses. A call center could begin with call type and time of day, then incorporate escalation chains and emotional load factors. The specific criteria will
vary by domain, but the overall methodology, including identifying impact factors, assigning ordinal weights, combining them multiplicatively rather than additively, and validating the measure against operational outcomes, remains applicable across settings.
All Hours Are Not the Same
The next phase of this work is to shift from units to individuals. Each call affects responders differently based on tenure, recovery, career exposure, and life experience. Biometric integration and AI-driven cognitive readiness assessments can help close that gap.
In operations research, assuming a utilization metric is perfectly accurate can be misleading. Operations relying on time-based utilization ultimately acknowledge that hours are not interchangeable. This study’s model assigns a specific value to reflect that distinction.
References
Brandt, J., Elliott, J., 2021, “Replacing UHU: Can a Scrappy Newcomer Topple This Long-reigning KPI?” Pinnacle Webinar Series, Fitch & Associates.
Ercolani, J., Cure, L., Misasi, P., 2024, “Relationship Between Conventional Workload Surrogates and VACP Assessments in Emergency Medical Services,” Proceedings of the IISE Annual Conference & Expo.
Fitch & Associates, 2017, “Fire Service Fatigue: A Problem You Can’t Afford to Ignore,” chrome-extension://efaidnbmnnnibpcajpcglclefindmkaj/https://fitchassoc.com/wp-content/uploads/2017/06/Fire-Service-Fatigue.pdf.
Marciniak, R. A., Cornell, D. J., Meyer, B. B., Azen, R., Laiosa, M. D., Ebersole, K. T., 2024, “Workloads of Emergency Call Types in Active-duty Firefighters,” Merits, Vol. 4, pp. 1-18.
Meyer, J. J., 2024, “Unit-Hour-Utilization Impact on EMS Provider Satisfaction,” U.S. Fire Administration, Executive Fire Officer Program applied research project.
Nardo, M., Saisana, M., Saltelli, A., Tarantola, S., Hoffmann, A., Giovannini, E., 2008, “Handbook on Constructing Composite Indicators: Methodology and User Guide,” OECD Publishing, https://publications.jrc.ec.europa.eu/repository/handle/JRC47008.
Saltelli, A., Ratto, M., Andres, T., Campolongo, F., Cariboni, J., Gatelli, D., Saisana, M., Tarantola, S., 2007, “Global Sensitivity Analysis. The Primer,” West Sussex, United Kingdom: John Wiley & Sons.
Wahl, C. A., Marciniak, R. A., Meyer, B. B., Ebersole, K. T., 2025, “Association Between Call Volume and Perceptions of Stress and Recovery in Active-duty Firefighters,” Fire, Vol. 8, No. 7, https://doi.org/10.3390/fire8070268.
Wendy Korotkin is the Data and Analytics Manager for Boulder Fire-Rescue. A former firefighter with 30 years in public safety, she works on fire and EMS data modeling, process improvement, planning, and performance measurement. She can be reached at [email protected]. Alex Stephenson is a career firefighter/paramedic for a Department of Fire & Rescue in the Northern Virginia area. He has served for 19 years and has a master’s degree in decision analytics from Virginia Commonwealth University. He uses data analytics to improve response performance, reallocate resources, and evaluate patient care. He can be reached at [email protected].
