This function estimates the average treatment effect (ATE, default) or average treatment effect on the treated (ATET, optional) of a continuously distributed treatment under selection on observables (unconfoundedness), using double machine learning with kernel smoothing around the treatment values of interest.
treatcontDML(
y,
d,
x,
dtreat,
dcontrol,
ATET = FALSE,
MLmethod = "lasso",
psmethod = 1,
trim = 0.1,
lognorm = FALSE,
bw = NULL,
bwfactor = 0.7,
cluster = NULL,
k = 3
)A list with the following components:
effect: Estimate of the average treatment effect (ATE) or average treatment effect on the treated (ATET).
se: Standard error of the effect estimate.
trimmed: Number of discarded (trimmed) observations.
pval: P-value.
Outcome variable. Should not contain missing values.
Treatment variable. Should be continuous and not contain missing values.
Covariates. Should not contain missing values.
Value of the treatment for which the mean potential outcome is evaluated as the "treatment" level.
Value of the treatment for which the mean potential outcome is evaluated as the "control" level.
Logical. If FALSE (default), the average treatment effect (ATE) in the total population is estimated. If TRUE, the average treatment effect on units with treatment close to dtreat (ATET) is estimated instead.
Machine learning method for estimating nuisance parameters using the SuperLearner package. Must be one of "lasso" (default), "randomforest", "xgboost", "svm", "ensemble", or "parametric".
Method for computing generalized propensity scores. Set to 1 for estimating conditional treatment densities using the treatment as dependent variable (imposing the parametric assumption of a Gaussian treatment distribution), or 2 for using the treatment kernel weights as dependent variable (nonparametric). Default is 1.
Trimming threshold (in percent) for discarding observations whose (normalized) weight would otherwise dominate the estimate. Default is 0.1.
Logical indicating if log-normal transformation should be applied when estimating conditional treatment densities using the treatment as dependent variable (parametric). Default is FALSE.
Bandwidth for kernel density estimation. Default is NULL, implying that the bandwidth is calculated based on the rule-of-thumb.
Factor by which the bandwidth is multiplied. Default is 0.7 (undersmoothing).
Optional clustering variable for calculating standard errors.
Number of folds in k-fold cross-fitting. Default is 3.
Under ATET=FALSE (default), the function estimates the ATE using a doubly robust, kernel-weighted score for continuous treatments, following Kennedy, Ma, McHugh, and Small (2017) and Colangelo and Lee (2020): for each of dtreat and dcontrol, every observation contributes both an outcome-regression prediction at that dose and a kernel/generalized-propensity-score-weighted residual correction. Under ATET=TRUE, a different doubly robust function instead makes observations near dcontrol to match the covariate distribution of units with treatment near dtreat, targeting the effect on that treated subpopulation rather than the total population.
Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., Robins, J. (2018): "Double/debiased machine learning for treatment and structural parameters", The Econometrics Journal, 21, C1-C68.
Kennedy, E. H., Ma, Z., McHugh, M. D., Small, D. S. (2017): "Non-parametric methods for doubly robust estimation of continuous treatment effects", Journal of the Royal Statistical Society Series B, 79, 1229-1245.
Colangelo, K., Lee, Y.-Y. (2020): "Double debiased machine learning nonparametric inference with continuous treatments", arXiv preprint 2004.03036.
if (FALSE) {
# Example with simulated data
n=3000
x=0.8*rnorm(n)
d=x+rnorm(n)
y=2*d+x+rnorm(n)
# true effect is 2
results=treatcontDML(y=y, d=d, dtreat=1, dcontrol=0, x=x, MLmethod="lasso")
cat("ATE: ", round(results$effect, 3), ", Standard error: ", round(results$se, 3))
}
Run the code above in your browser using DataLab