Adaptive Experimental Design for Policy Learning
Working paper.
- policy learning
- contextual best-arm identification
- simple regret
- minimax rate
In one sentence. When the deliverable of an experiment is a policy rather than a single recommended arm, the natural criterion is worst-case expected simple regret; the proposed PLAS strategy is minimax rate-optimal for it.
From best arm to best policy
Classical best-arm identification asks which single treatment is best on average. Many decisions are not of that form: the best treatment depends on the context, and what the decision-maker wants out of the experiment is a policy — a function mapping observed covariates to a recommended arm.
This paper studies the contextual best-arm identification problem in that light. During the experiment the decision-maker assigns arms to units; at the end, the decision-maker outputs a policy. Performance is measured by worst-case expected regret: the gap between the expected outcome of an optimal policy and that of the recommended one.
Results
- A lower bound for the expected simple regret in this setting.
- Adaptive Sampling-Policy Learning (PLAS), a strategy that interleaves adaptive sampling with policy learning.
- PLAS is minimax rate-optimal: the leading factor of its regret upper bound matches the lower bound as the number of experimental units grows.
Why it matters
Experiments that are run to inform targeting decisions — who should receive a subsidy, which users should see which intervention — are evaluated on the quality of the resulting rule, not on the precision of an average effect. Designing the experiment against that criterion, rather than against estimation variance, changes what the optimal allocation looks like.