Small-Loss Bounds for Online Learning with Partial Information

Thodoris Lykouris
Thodoris Lykouris
[email protected]
https://orcid.org/0000-0002-3375-5579
Massachusetts Institute of Technology, Cambridge, Massachusetts 02139;
Search for more papers by this author
,
Karthik Sridharan
Karthik Sridharan
[email protected]
Cornell University, Ithaca, New York 14850
Search for more papers by this author
,
Éva Tardos
Éva Tardos
[email protected]
Cornell University, Ithaca, New York 14850
Search for more papers by this author

Massachusetts Institute of Technology, Cambridge, Massachusetts 02139;

Search for more papers by this author

Karthik Sridharan

[email protected]

Cornell University, Ithaca, New York 14850

Search for more papers by this author

Éva Tardos

[email protected]

Cornell University, Ithaca, New York 14850

Search for more papers by this author

Published Online:25 Jan 2022https://doi.org/10.1287/moor.2021.1204

Abstract

We consider the problem of adversarial (nonstochastic) online learning with partial-information feedback, in which, at each round, a decision maker selects an action from a finite set of alternatives. We develop a black-box approach for such problems in which the learner observes as feedback only losses of a subset of the actions that includes the selected action. When losses of actions are nonnegative, under the graph-based feedback model introduced by Mannor and Shamir, we offer algorithms that attain the so called “small-loss” $o (α L^{⋆})$ regret bounds with high probability, where α is the independence number of the graph and $L^{⋆}$ is the loss of the best action. Prior to our work, there was no data-dependent guarantee for general feedback graphs even for pseudo-regret (without dependence on the number of actions, i.e., utilizing the increased information feedback). Taking advantage of the black-box nature of our technique, we extend our results to many other applications, such as combinatorial semi-bandits (including routing in networks), contextual bandits (even with an infinite comparator class), and learning with slowly changing (shifting) comparators. In the special case of multi-armed bandit and combinatorial semi-bandit problems, we provide optimal small-loss, high-probability regret guarantees of $\tilde{O} (\sqrt{d L^{⋆}})$ , where d is the number of actions, answering open questions of Neu. Previous bounds for multi-armed bandits and semi-bandits were known only for pseudo-regret and only in expectation. We also offer an optimal $\tilde{O} (\sqrt{κ L^{⋆}})$ regret guarantee for fixed feedback graphs with clique-partition number at most κ.

cover image Mathematics of Operations Research

Volume 47, Issue 3

August 2022

Pages 1707-2545, C2

Article Information

Metrics

Information

Received:June 13, 2018
Accepted:June 30, 2021
Published Online:January 25, 2022

Cite as

Thodoris Lykouris, Karthik Sridharan, Éva Tardos (2022) Small-Loss Bounds for Online Learning with Partial Information. Mathematics of Operations Research 47(3):2186-2218.

https://doi.org/10.1287/moor.2021.1204

Keywords

PDF download

Available Issues

Available Issues

Available Issues

Available Issues

Available Issues

Available Issues

Available Issues

Small-Loss Bounds for Online Learning with Partial Information

Abstract

Volume 47, Issue 3

Article Information

Metrics

Information

Cite as

Keywords

Sign Up for INFORMS Publications Updates and News