Active Adaptive Experimental Design for Treatment Effect Estimation with Covariate Choices
International Conference on Machine Learning (ICML), 2024. Oral presentation (top 1.5%).
- adaptive experiments
- semiparametric efficiency
- ATE
- covariate design
In one sentence. If the experimenter can choose who enters the experiment as well as what they receive, then jointly optimizing the covariate density and the propensity score reaches a strictly lower semiparametric efficiency bound than optimizing the propensity score alone.
Setting
In each round of the experiment the experimenter samples a unit, assigns a treatment, and immediately observes the outcome. At the end of the experiment the average treatment effect is estimated from the accumulated sample, and the design is judged by the asymptotic variance of that estimate.
The existing literature on adaptive experimental design treats the propensity score — the treatment-assignment probability — as the design variable, and adapts it as the experiment reveals which arms are noisier. This paper adds a second design variable: the covariate density, i.e. which covariate profiles the experimenter chooses to enroll in the first place.
What the paper does
- Derives the efficient covariate density and propensity score that minimize the semiparametric efficiency bound, and shows that optimizing both lowers the bound more than optimizing the propensity score alone.
- Designs an adaptive experiment that estimates these two quantities sequentially from the data collected so far, and uses them to decide whom to sample next and which treatment to assign.
- Proposes an ATE estimator whose asymptotic variance coincides with the minimized bound.
When this design is the right tool
Two conditions have to hold: the experimenter can decide whether to include a unit after observing its covariates, and outcomes are observed quickly enough to inform later rounds. Where both hold and each observation is expensive — advertising trials, or clinical studies where enrollment is designed rather than merely accepted — the efficiency gain translates directly into a smaller experiment.