Learn R Programming

causalweight (version 1.1.6)

catehetDML: Testing effect homogeneity across studies with double machine learning

Description

Testing effect homogeneity across studies with double machine learning

Usage

catehetDML(
  y,
  d,
  x,
  z,
  MLmethod = "lasso",
  k = 5,
  L = 4,
  trim = 0.05,
  normalized = TRUE,
  seed = 123
)

Value

A list with the following elements:

teststat: estimate of the target parameter, which is zero under the null hypothesis of effect homogeneity across studies.

se: standard error of teststat.

pval: p-value of the two-sided test of the null hypothesis, based on the t-statistic teststat/se.

n_eff: effective sample size after trimming and inverse probability weighting.

n_eff_z: effective sample size for each study-specific contrast.

ntrimmed: number of observations discarded by trimming.

studies: labels of the studies for which contrasts are formed.

psi: individual contributions to the score function.

w_full: inverse probability-based weights, zero for trimmed observations.

Arguments

y

Dependent variable, must not contain missings.

d

Treatment variable, must be binary (0/1), must not contain missings.

x

Covariates, must not contain missings.

z

Study indicator across which effect homogeneity is tested, must not contain missings. If z is a factor, a character vector, or a numeric vector taking only integer values, each distinct value defines one study and contrasts are formed for all but the last of them. If z is numeric and not integer-valued, it is discretised into L bins of equal size, each of which is then treated as one study.

MLmethod

Machine learning method for estimating the nuisance parameters based on the SuperLearner package. Must be either "lasso" (default) for lasso estimation, "randomforest" for random forests, "xgboost" for xg boosting, "svm" for support vector machines, "ensemble" for using an ensemble algorithm based on all previously mentioned machine learners, or "parametric" for linear or logit regression.

k

Number of folds in k-fold cross-fitting. Default is 5.

L

Number of bins into which z is discretised if z is numeric and not integer-valued. Ignored otherwise, in which case the number of studies is determined by the distinct values of z. Default is 4.

trim

Trimming rule for discarding observations whose estimated joint probabilities of receiving a treatment state and belonging to a study given x are smaller than trim or larger than 1-trim (to avoid too small denominators in weighting by the inverse of these probabilities). Default is 0.05.

normalized

If set to TRUE, then the inverse probability-based weights are normalized such that they average to one within each combination of treatment state and study. Default is TRUE.

seed

Default is 123.

Details

Tests whether the conditional average treatment effect (CATE) of a binary treatment d on an outcome y given covariates x is homogeneous across the studies indexed by z. The null hypothesis is that for each study, the CATE within that study coincides with the CATE in the remaining studies at any value of x. If all studies are randomized experiments, a rejection indicates that treatment effects vary with characteristics that differ across studies but are not contained in x, which speaks to the external validity of the experimental effects. If some studies are experimental while others are observational, a rejection may in addition reflect unobserved confounding in the observational studies, such that differences between the estimands may arise from unobserved confounding, effect heterogeneity, or both. See Armendariz and Huber (2026). Estimation is based on a Neyman-orthogonal score that is quadratic in the doubly robust double difference of conditional mean outcomes across treatment states and studies, in the spirit of Apfel, Hatamyar, Huber and Kueck (2024), combined with k-fold cross-fitting. The nuisance parameters, namely the conditional mean outcomes \(E[Y|D=d,X,Z=z]\) and the joint probabilities \(Pr(D=d,Z=z|X)\), are estimated by the machine learner selected in MLmethod. The test statistic is asymptotically normal under the null hypothesis.

References

Armendariz, A., Huber, M. (2026): Testing Effect Homogeneity and Confounding in High-Dimensional Experimental and Observational Studies. arXiv:2602.19703.

Apfel, N., Hatamyar, J., Huber, M., Kueck, J. (2024): Learning control variables and instruments for causal analysis in observational data. arXiv:2407.04448.

Examples

Run this code
if (FALSE) {
# Two studies with homogeneous CATEs, such that the test should not reject.
set.seed(123)
n <- 2000
x <- matrix(rnorm(n * 5), ncol = 5)
z <- sample(1:2, n, replace = TRUE)                     # two studies
d <- rbinom(n, 1, plogis(0.5 * x[, 1] + 0.5 * (z == 2)))
y <- d + x[, 2] + rnorm(n)                              # CATE is 1 in both studies
output <- catehetDML(y = y, d = d, x = x, z = z)
output$teststat
output$pval

# Same design, but with CATEs differing across studies, such that the test should reject.
y <- d + 1 * (z == 2) * d + x[, 2] + rnorm(n)           # CATE is 1 vs. 2
output <- catehetDML(y = y, d = d, x = x, z = z)
output$teststat
output$pval
}

Run the code above in your browser using DataLab