Lexicographic Multiarmed Bandit

07/26/2019

∙

We consider a multiobjective multiarmed bandit problem with lexicographically ordered objectives. In this problem, the goal of the learner is to select arms that are lexicographic optimal as much as possible without knowing the arm reward distributions beforehand. We capture this goal by defining a multidimensional form of regret that measures the loss of the learner due to not selecting lexicographic optimal arms, and then, consider two settings where the learner has prior information on the expected arm rewards. In the first setting, the learner only knows for each objective the lexicographic optimal expected reward. In the second setting, it only knows for each objective near-lexicographic optimal expected rewards. For both settings we prove that the learner achieves expected regret uniformly bounded in time. In addition, we also consider the harder prior-free case, and show that the learner can still achieve sublinear in time gap-free regret. Finally, we experimentally evaluate performance of the proposed algorithms in a variety of multiobjective learning problems.

READ FULL TEXT

Lexicographic Multiarmed Bandit

On Regret-Optimal Learning in Decentralized Multi-player Multi-armed Bandits

Combinatorial Bandits without Total Order for Arms

Multi-objective Contextual Bandit Problem with Similarity Information

Repeated A/B Testing

Safe Linear Stochastic Bandits

Compliance-Aware Bandits

Preselection Bandits under the Plackett-Luce Model

Lexicographic Multiarmed Bandit

Related Research

On Regret-Optimal Learning in Decentralized Multi-player Multi-armed Bandits

Combinatorial Bandits without Total Order for Arms

Multi-objective Contextual Bandit Problem with Similarity Information

Repeated A/B Testing

Safe Linear Stochastic Bandits

Compliance-Aware Bandits

Preselection Bandits under the Plackett-Luce Model