Dynamic Contract Design with Learning
Abstract
We investigate a dynamic moral hazard problem in which the arrival rate of customers is unknown to both the principal and the agent in the beginning and can be learned over time. Specifically, the agent can exert effort to generate random arrivals that are beneficial to the principal. Neither the principal nor the agent knows the arrival rate, but they can learn more about it over time from observing arrivals. Not knowing the arrival rate, as well as whether the agent exerts effort, makes it very challenging to identify the optimal dynamic contract. Therefore, we focus on designing dynamic contracts that achieve the best regret rate. We consider a discrete-time setting with two possible market conditions, one favorable and the other not, and propose two types of contracts, both of which achieve the regret rate of for a time horizon T. Our contracts are easy to implement for the principal. They motivate the agent to always exert effort before termination and therefore are easy for the agent to respond to as well. We also establish a regret-rate lower bound, , among all contracts, including ones that allow the agent to shirk. This implies that our regret rate is the best possible. We further extend these results to continuous-time settings and settings where the principal and the agent have different prior beliefs.
Funding: Y. Liang is supported by the National Natural Science Foundation of China [Grant 72325001].
Supplemental Material: All supplemental materials, including the code, data, and files required to reproduce the results, are available at https://doi.org/10.1287/opre.2024.0944.

