Estimation of natural direct and indirect quantile treatment effects (QTEs) of a binary treatment operating through a scalar mediator, based on double machine learning. The method inverts post-lasso-based double/debiased machine learning estimates of the cumulative distribution functions (cdfs) of the four potential outcomes Y(1,M(1)), Y(1,M(0)), Y(0,M(1)), Y(0,M(0)), computed from the efficient influence functions of Hsu, Huber, and Yen (2026), and provides pointwise and uniform confidence bands via a multiplier bootstrap.
medqteDML(
y,
d,
m,
x,
tau = seq(0.2, 0.9, 0.1),
a = NULL,
kfold = 5,
trim = 0.05,
nboot = 1000,
alpha = 0.05,
q = 0.1
)A medqteDML object contains the following components:
effects: a list of five data frames, one per estimand: total (the total effect, Y(1,M(1)) minus Y(0,M(0))); dir.treat and dir.control (the direct effects with the mediator fixed at its value under treatment, M(1), or under non-treatment, M(0), respectively -- NDQTE'} and \code{NDQTE} in Hsu, Huber, and Yen 2026); and \code{indir.treat} and \code{indir.control} (the indirect effects with the treatment fixed at \code{D=1} or \code{D=0}, respectively -- \code{NIQTE} and \code{NIQTE' in Hsu, Huber, and Yen 2026). Each data frame has columns tau (the quantile rank), effect (the point estimate), se (bootstrap standard deviation), pw.lower/pw.upper (pointwise confidence interval), and unif.lower/unif.upper (uniform confidence band).
cdf: a data frame with the estimated cdfs of the four potential outcomes across the grid a.
quantiles: a data frame with the estimated quantiles Q(1,M(1)), Q(1,M(0)), Q(0,M(1)), Q(0,M(0)) at each tau.
boot: a list with the raw multiplier bootstrap draws of the five QTEs (rows are tau, columns are bootstrap replications), for users who want to construct alternative confidence measures.
ntrimmed: number of observations whose estimated propensity score P(D=1|X) or P(D=1|M,X) was winsorized (capped to trim or 1-trim) in at least one fold or grid point.
Outcome/dependent variable, must be a numeric scalar variable, must not contain missings.
Treatment, must be binary (either 1 or 0), must not contain missings.
Mediator, must be a numeric scalar variable, must not contain missings.
(Potential) pre-treatment confounders of the treatment, mediator, and/or outcome, must not contain missings.
Vector of quantile ranks at which the QTEs are estimated, each strictly between 0 and 1. Default is seq(0.2,0.9,0.1).
Grid of values at which the cdfs of the potential outcomes are estimated before inversion into quantiles. Default is NULL, in which case a grid of at least 50 points is built from empirical quantiles of y, spanning somewhat beyond the range of tau: the potential-outcome distributions are shifted relative to the pooled distribution of y, so a grid tied only to tau itself is usually too narrow and forces the quantile inversion to extrapolate. A custom a can be supplied for more control, but should be finer and wider than tau; medqteDML warns if the estimated cdfs do not bracket the requested tau.
Number of folds in the cross-fitting procedure. Default is 5.
Trimming threshold applied to the estimated propensity scores P(D=1|X) and P(D=1|M,X), which are winsorized to [trim, 1-trim]. Default is 0.05.
Number of multiplier bootstrap replications used for the pointwise confidence intervals and uniform confidence bands. Default is 1000. The bootstrap reuses the estimated efficient influence functions rather than refitting the nuisance parameters, so increasing nboot is inexpensive relative to increasing length(a) or kfold.
Significance level for the confidence intervals and bands. Default is 0.05.
Lower quantile used to compute the rescaled quantile spread underlying the uniform confidence bands, see Chernozhukov, Fernandez-Val, and Melly (2013). Default is 0.1, a common choice in that literature.
The nuisance parameters (propensity scores P(D=1|X) and P(D=1|M,X), and the conditional cdf P(Y<=a|D,M,X) estimated by distribution regression) are estimated by post-lasso, with covariates selected via a double/triple selection step (lasso selection from the D|X, D|M,X, and 1{Y<=a}|D,M,X models, unioned) at every grid point a and every fold, following the implementation underlying the empirical application in Hsu, Huber, and Yen (2026). The estimator itself is the K-fold cross-fitting estimator of the cdf of each potential outcome based on the (triply robust) efficient influence function, see Section 2 of Hsu, Huber, and Yen (2026) and, for the underlying regression imputation approach, Algorithm 2 in Farbmacher et al. (2022). Quantiles are obtained by inverting the estimated cdfs on the grid a, after a rearrangement step (Chernozhukov, Fernandez-Val, and Galichon 2010) that enforces monotonicity. Statistical inference is based on a multiplier bootstrap that reuses the estimated efficient influence functions, following Section 2.5 of Hsu, Huber, and Yen (2026); both pointwise confidence intervals (percentile/reflection method) and uniform confidence bands (based on the bootstrap Kolmogorov-Smirnov max-t-statistic and the rescaled quantile spread) are provided.
Chernozhukov, V., Fernandez-Val, I., and Galichon, A. (2010): "Quantile and Probability Curves Without Crossing", Econometrica, 78, 1093-1125.
Chernozhukov, V., Fernandez-Val, I., and Melly, B. (2013): "Inference on Counterfactual Distributions", Econometrica, 81, 2205-2268.
Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., Robins, J. (2018): "Double/debiased machine learning for treatment and structural parameters", The Econometrics Journal, 21, C1-C68.
Farbmacher, H., Huber, M., Laffers, L., Langen, H., and Spindler, M. (2022): "Causal mediation analysis with double machine learning", The Econometrics Journal, 25, 277-300.
Hsu, Y.-C., Huber, M., and Yen, Y.-M. (2026): "Estimation of Direct and Indirect Quantile Treatment Effects with Double Machine Learning", Journal of Business & Economic Statistics, DOI: 10.1080/07350015.2026.2654889.
# A little example with simulated data
if (FALSE) {
n=5000 # sample size
p=10 # number of covariates
x=matrix(rnorm(n*p),ncol=p) # covariate matrix
d=rbinom(n,1,plogis(0.5*x[,1]-0.3*x[,2])) # treatment equation
m=rbinom(n,1,plogis(-0.3+0.8*d+0.4*x[,1]-0.3*x[,2])) # mediator equation
y=2*d+1.5*m-0.4*x[,1]+0.4*x[,2]+rnorm(n) # outcome equation
# Direct effect is 2 at every quantile; indirect effect is around 0.3
output=medqteDML(y=y,d=d,m=m,x=x,tau=c(0.3,0.5,0.7))
output$effects$dir.control
output$effects$indir.treat
output$effects$dir.treat
output$effects$indir.control
output$ntrimmed
}
Run the code above in your browser using DataLab