Service Systems with On-Demand and Reserved Servers

Published Online:https://doi.org/10.1287/msom.2024.1504

Problem definition: In many real-world applications, such as cloud computing, there is a growing trend to supplement long-term reserved capacity with short-term on-demand capacity. We study a queueing system that employs both reserved and (relatively more expensive) on-demand servers. The number of reserved servers is decided at the beginning of the time horizon, while the number of on-demand servers is decided dynamically in real time. The objective is to minimize the costs incurred in hiring servers and in job waiting.

Methodology/results: Using a sample-path analysis, we establish that the optimal on-demand control is a threshold-based bang-bang policy: If the number of jobs in the system is below a threshold, no on-demand servers are employed. Otherwise, the number of on-demand servers is chosen such that no jobs wait. We present algorithms for obtaining an ϵ-optimal solution for the discounted-cost problem and an optimal solution for the average-cost problem. Furthermore, to address operational frictions, we extend our analysis to incorporate fixed switching costs and provisioning latency. We propose a periodic tracking policy and prove its asymptotic optimality.

Managerial implications: Our analysis provides practical guidance for managing capacity and waiting costs in queueing systems that have access to both long-term and short-term server capacity. Our numerical experiments provide prescriptive insights on tailoring reserved and on-demand capacity to job urgency and volume.

INFORMS site uses cookies to store information on your computer. Some are essential to make our site work; Others help us improve the user experience. By using this site, you consent to the placement of these cookies. Please read our Privacy Statement to learn more.