Learn R Programming

causalweight (version 1.1.6)

didmedDMLpanel: Difference-in-Differences for Mediation Analysis with Panel Data and Discrete Treatments Using Double Machine Learning

Description

This function estimates the total effect, natural direct effect, and natural indirect effect for the treated group in the post-treatment period in panel data with discrete treatments. Estimation is based on a difference-in-differences approach for mediation analysis combined with double machine learning to control for (possibly time-varying) confounders in a data-driven manner. The function supports various machine learning methods for estimating nuisance parameters through k-fold cross-fitting.

Usage

didmedDMLpanel(
  y0,
  y1,
  d,
  m,
  x,
  dtreat = 1,
  dcontrol = 0,
  MLmethod = "lasso",
  trim = 0.05,
  cluster = NULL,
  k = 3
)

Value

A list with the following components:

eff: Estimates of the total effect, natural direct effect, and natural indirect effect for the treated group in the post-treatment period.

se: Standard error of the estimates.

tval: t-value of the estimates.

pval: p-value of the estimates.

trimmed: Number of discarded (trimmed) observations.

Arguments

y0

Outcome variable in the pre-treatment period. Should not contain missing values.

y1

Outcome variable in the post-treatment period. Should not contain missing values.

d

Treatment group indicator (discrete). Should not contain missing values.

m

Mediator variable. Should not contain missing values.

x

Covariates to be controlled for. Should not contain missing values.

dtreat

Value of the treatment under treatment (in the treatment period of interest). Default is 1.

dcontrol

Value of the treatment under control (in the treatment period of interest). Default is 0.

MLmethod

Machine learning method for estimating nuisance parameters using the SuperLearner package. Must be one of "lasso" (default), "randomforest", "xgboost", "svm", "ensemble", or "parametric".

trim

Trimming threshold for discarding observations with too small propensity scores in the control group. Default is 0.05.

cluster

Optional clustering variable for calculating cluster-robust standard errors.

k

Number of folds in k-fold cross-fitting. Default is 3.

Details

This function estimates the total effect, natural direct effect, and natural indirect effect for the treated group in the post-treatment period in panel data with discrete treatments. Estimation is based on the difference-in-differences approach for mediation analysis proposed by Huber and Oberhänsli (2026). Specifically, these effects are computed from normalized sample analogs of the doubly robust expressions in equations (27) and (28). Double machine learning is used to control for confounders in a data-adaptive way. The function supports different machine learning methods for estimating nuisance parameters (conditional mean outcomes and propensity scores) as well as cross-fitting to mitigate overfitting.

References

Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., Robins, J. (2018): "Double/debiased machine learning for treatment and structural parameters", The Econometrics Journal, 21, C1-C68.

Huber, M., and Oberhänsli, S. J. (2026): "Difference-in-differences for mediation analysis using double machine learning", arXiv preprint 2602.23877.

Examples

Run this code
if (FALSE) {
# Example with simulated data
n=4000                            # sample size
u=rnorm(n)                        # time constant unobservable
x=rnorm(n)                        # covariate
d=1*(x+0.5*u+rnorm(n)>0)          # treatment
m=x+0.5*d+rnorm(n)                # mediator
y0=u+rnorm(n)                     # outcome in the pre-treatment period
y1=x+1+d+m+u+rnorm(n)             # outcome in the post-treatment period
# true NDET is equal to 1; true NIET is equal to 0.5; true ATET is equal to 1.5
didmedDMLpanel(y0=y0,y1=y1, d=d, m=m, x=x)
}

Run the code above in your browser using DataLab