Learn R Programming

causalweight (version 1.1.6)

detectIV: Detection of instruments and control variables and validity testing with double machine learning

Description

Tests the validity of a pre-specified instrument and pre-specified control variables (mode 1) or learns the partition of instruments and control variables from the data (mode 2), following the two-step procedure of Apfel, Hatamyar, Huber, and Kueck (2025).

Mode 1 — user-specified x and z: The user provides the covariates x and a single candidate instrument z directly. The function skips instrument selection and runs only the validity test: it tests whether z is conditionally independent of y given d and x.

Mode 2 — data-driven selection via q: The user provides a matrix q of candidate variables without specifying which are instruments and which are controls. The function runs the full two-stage procedure: (1) a first-stage DML relevance test selects candidate instruments associated with d, and (2) the validity test retains those that are conditionally independent of y.

Usage

detectIV(
  y,
  d,
  z = NULL,
  x = NULL,
  q = NULL,
  alpha = 0.1/log(length(y)),
  critval = 0.3,
  MLmethod = "lasso",
  k = 3,
  trim = 0.01,
  L = 4,
  seed = 123
)

Value

Mode 1 returns a list with:

teststat

Test statistic (should be near zero under the null).

se

Standard error of the test statistic.

pval

P-value of the exclusion-restriction test.

n_eff

Number of observations retained after trimming.

critval

Significance level used for the exclusion test.

Mode 2 returns a list with:

valid_IVs

Integer vector of column indices of q detected as valid IVs. Empty if none found.

firststage_pass

Column indices passing the first-stage test (\(\hat{S}\)).

firststage_pvals

Named numeric vector of first-stage p-values for all columns of q.

exclusion_pvals

Named numeric vector of exclusion-restriction p-values. NA for columns not in \(\hat{S}\).

exclusion_results

Named list of full test output objects for each first-stage candidate.

alpha

Significance level used for the first-stage test.

critval

Significance level used for the exclusion test.

Arguments

y

Outcome variable, numeric vector. Must not contain missings.

d

Treatment variable, numeric vector. Must not contain missings.

z

Mode 1 only. Candidate instrument, numeric vector. Must not contain missings. Provide together with x; leave NULL (default) to use mode 2 with q.

x

Mode 1 only. Covariate matrix. Must not contain missings. Provide together with z; leave NULL (default) to use mode 2 with q.

q

Mode 2 only. Matrix of candidate variables whose columns are iteratively considered as potential instruments (with the remaining columns serving as control variables). Must not contain missings and must have at least two columns. Leave NULL (default) to use mode 1 with x and z.

alpha

Mode 2 only. Significance level for the first-stage relevance test. A candidate passes if its DML p-value is below alpha. The paper recommends 0.1 / log(n). Default is 0.1 / log(length(y)).

critval

Significance level for the exclusion-restriction test. Candidates whose p-value exceeds critval are classified as valid IVs (null of conditional independence not rejected). Note that a higher value is a more conservative choice; Apfel et al. recommend 0.3. Default is 0.3.

MLmethod

Machine learning method for estimating nuisance parameters via the SuperLearner package. Must be one of "lasso" (default), "randomforest", "xgboost", "svm", "ensemble", or "parametric".

k

Number of folds in k-fold cross-fitting. Default is 3.

trim

Trimming threshold for propensity scores: observations whose estimated propensity score falls below trim or above 1 - trim are discarded. Default is 0.01.

L

Number of partition bins for non-binary instruments with 15 or more unique values. Variables with fewer than 15 unique values use their natural categories. Default is 4.

seed

Random seed. Default is 123.

Details

Mode 1 runs the validity test of Apfel et al. (2025) directly on the user-supplied z and x, testing \(H_0: E[Y \mid D, X] = E[Y \mid D, X, Z]\). This is appropriate when the researcher has a single pre-specified candidate IV and wants to assess its validity (and that of the control variables) without selection.

Mode 2 implements the full data-driven algorithm of Apfel et al. (2025). For a matrix \(Q\) of \(p\) candidate variables it constructs partitions \(P_j = \{Z = Q_j,\, X = Q_{[j]}\}\) and:

  1. Fits a DML partially linear regression (PLR) of d on each column \(Q_j\) controlling for \(Q_{[j]}\), retaining columns whose p-value is below alpha as first-stage candidates \(\hat{S}\).

  2. Applies the validity test to each candidate in \(\hat{S}\), classifying those with p-value above critval as valid IVs \(\hat{V}\).

The detected set is \(\hat{P}_{\text{pass}} = \hat{S} \cap \hat{V}\).

In both modes the validity test uses the doubly robust score of Apfel et al. (2025): eq. (7) for binary instruments and the multi-partition score of eq. (18) / Appendix C for non-binary instruments.

References

Apfel, N., Hatamyar, J., Huber, M., Kueck, J. (2025): "Learning control variables and instruments for causal analysis in observational data," arXiv:2407.04448.

Huber, M., Kloiber, K., Lafférs, L. (2026): "Testing Full Mediation of Treatment Effects and the Identifiability of Causal Mechanisms," arXiv:2603.04109.

Examples

Run this code
if (FALSE) {
set.seed(42)
n <- 2000; p <- 10
Sigma <- outer(1:p, 1:p, function(i, j) 0.5^abs(i - j))
Q_raw <- mvtnorm::rmvnorm(n, rep(0, p), Sigma)
pis   <- 1 / (1 + exp(-2 * Q_raw))
Q     <- matrix(rbinom(n * p, 1, pis), nrow = n)
beta  <- c(0.8 / (1:4), rep(0, 5))
W <- rnorm(n); U <- rnorm(n); V <- rnorm(n)
D <- as.numeric(Q[, 1:9] %*% beta + Q[, p] + V > 0)
Y <- D + Q[, 1:9] %*% beta + W + U   # Q[, p] has no direct effect on Y

# Mode 1: test a pre-specified instrument (last column) directly
res1 <- detectIV(y = Y, d = D, z = Q[, p], x = Q[, 1:9])
cat("Mode 1 p-value:", round(res1$pval, 3), "\n")

# Mode 2: learn instruments and controls from Q
res2 <- detectIV(y = Y, d = D, q = Q, alpha = 0.1 / log(n))
cat("Mode 2 column indices of detected valid IVs:", res2$valid_IVs, "\n")
}

Run the code above in your browser using DataLab