Fits the stochastic frontier model of Amsler, Prokhorov and Schmidt (2016) in
which one or more regressors is endogenous, in the simultaneous-equation sense:
correlated with the statistical noise \(v\). Three estimators are provided --
full information maximum likelihood, a two-step control function, and corrected
2SLS. With uhet the model becomes that of Amsler, Prokhorov and Schmidt
(2017), in which environmental variables entering the inefficiency scale may
themselves be endogenous.
ivsfm(formula, endogenous, instruments, data,
model_name = c("IVLIML", "IVCF", "C2SLS"), uhet = NULL,
inefdec = TRUE, maxit.bobyqa = 10000, maxit.psoptim = 1000,
maxit.optim = 1000, start_val = FALSE, PSopt = FALSE,
optHessian = TRUE, Method = "L-BFGS-B", verbose = FALSE,
rand.psoptim = NULL)An object of class "sfareg". out is a p x 3 matrix, one
row per parameter, columns par, st_err and t-val.
Parameters are the frontier coefficients, sigma_u, sigma_v, any
delta_ terms from uhet, and one rho_ per endogenous
variable (absent for "C2SLS"). Alongside the usual components:
Reduced-form coefficient matrix.
Reduced-form error covariance.
Correlation of each reduced-form error with \(v\).
The conditional noise standard deviation \(\sigma_c\), which
is smaller than sigma_v whenever \(\rho \ne 0\).
The 2SLS coefficients, which seed the likelihood and are the
estimate "C2SLS" corrects.
Whether the 2SLS residuals are skewed the wrong way.
\(E[u_i \mid \varepsilon_i, \xi_i]\) and
exp(-jlms).
The full \(q \times q\) covariance of \(\hat\rho\), which
endogeneity_test needs for the joint Wald statistic.
Standard errors for \(\rho\) arrived in 1.2.0. They were previously
reported as NA, on the ground that \(\rho = t/\sqrt{1+t't}\) has a
non-diagonal Jacobian and a delta-method value computed as though it were
diagonal would be wrong. That much was true, but the Jacobian itself is short --
\(J = (I - \rho\rho')/s\) with \(s = \sqrt{1+t't}\) -- so
\(\mathrm{Var}(\hat\rho) = J\,\mathrm{Var}(\hat t)\,J'\) is available
exactly. Checked against a numerical Jacobian in the tests, and validated
end-to-end by the size of endogeneity_test, which is 0.046 at a
nominal 5% over 1000 replications.
The structural frontier, y ~ x1 + x2, listing all
regressors including the endogenous ones.
One-sided formula naming the endogenous variables,
~ x2. Each must appear in formula or in uhet.
One-sided formula of the excluded instruments,
~ w1 + w2. The included exogenous regressors are added automatically,
so only the outside instruments are listed here. There must be at least as
many as there are endogenous variables.
A data.frame holding every variable in the four formulas.
"IVLIML" (the default), "IVCF" or
"C2SLS"; see ‘Details’.
Optional one-sided formula of environmental variables, giving
\(\sigma_{U,i} = \sigma_U \exp(q_i'\delta)\). Supplying it selects the
2017 model; omitting it gives the 2016 one. Variables may appear in both
uhet and endogenous.
TRUE (default) for a production frontier, where
inefficiency is subtracted; FALSE for a cost frontier.
Iteration caps for the three
optimizer stages. Ignored by "C2SLS", which is closed-form.
Name the starting-value vector in the returned object.
Run the particle-swarm stage between BOBYQA and optim.
Compute the Hessian, and with it the standard errors.
Method passed to optim for the final stage.
Report optimizer progress.
Optional seed for the particle-swarm stage.
Unlike the rest of the package neither formula nor endogenous nor
instruments takes a | segment. A pipe means “variance
determinant” in sfm and psfm, and reusing it here
for ivreg-style instrument syntax would give the same character two
meanings across entry points, so a pipe is an error rather than being
reinterpreted.
The model. Writing \(p_i\) for the endogenous block and \(z_i\) for the instruments, $$y_i = \beta'x_i + v_i - u_i, \qquad p_i = \Pi'z_i + \xi_i$$ $$u_i = \sigma_{U,i}|U_i|, \quad U_i \sim N(0,1), \qquad (v_i, \xi_i') \sim N(0, \Omega)$$ with \(u_i\) independent of \((v_i, \xi_i)\). Endogeneity is the correlation between \(\xi\) and \(v\): the regressors are correlated with the noise, not with the inefficiency.
Conditioning \(v\) on \(\xi\) turns this into an ordinary normal--half normal frontier in a shifted residual with a smaller noise variance, $$\mu_{c,i} = \Sigma_{v\xi}\Sigma_{\xi\xi}^{-1}\xi_i, \qquad \sigma_c^2 = \sigma_V^2 - \Sigma_{v\xi}\Sigma_{\xi\xi}^{-1}\Sigma_{\xi v}$$ which is what makes the likelihood closed-form and needs no simulation.
The estimators.
"IVLIML"APS (2016) Equation (13): maximum likelihood over the frontier and the reduced form jointly, so \(\Pi\) and \(\Sigma_{\xi\xi}\) are estimated alongside \(\beta\). Efficient, and its standard errors are correct as reported.
"IVCF"The two-step control function of Kutlu (2010), described
in APS (2016) Section 4.4. \(\Pi\) and \(\Sigma_{\xi\xi}\) are fixed at
their reduced-form least-squares values and only the frontier block is
maximized. Cheaper, still consistent, but it discards the information about
the reduced form carried in the frontier likelihood; its conventional
standard errors understate the true ones unless \(\Sigma_{v\xi} = 0\).
Measured over 30 replications at \(n = 1000\): with \(\rho = 0.6\)
the mean reported standard error on the endogenous coefficient is 0.912 of
the Monte Carlo standard deviation, while "IVLIML"'s is 1.005; with
\(\rho = 0\) the two are identical, as the theory says they should be.
Bootstrap, or use "IVLIML", when the standard errors matter.
"C2SLS"APS (2016) Section 4.1: 2SLS, then \(\sigma_U\) and
\(\sigma_V\) from the second and third moments of the residuals, then the
intercept corrected by \(\sqrt{2/\pi}\,\hat\sigma_U\). The residuals use
the actual endogenous regressors, not their fitted values, which the
paper flags explicitly. Moment-based rather than likelihood-based, so the
fit carries no $opt and logLik returns NA with
a warning -- the same footing as psfm's "GTRE_SEQ1". It
estimates no \(\rho\), since it never models the correlation it corrects
for.
Inefficiency. jlms implements APS (2016) Equations (14)--(15),
which condition on \(\xi\) as well as on \(\varepsilon\). Since \(\xi\)
is correlated with \(v\) it carries information about \(u\) even though
\(u\) is independent of \(\xi\), so this is a strictly better predictor
than the plain Jondrow et al. (1982) formula.
What is assumed, and what is not covered. \(u\) is assumed independent of \((v, \xi)\). The case where the reduced-form error is correlated with the inefficiency as well requires the copula likelihood of APS (2016) Section 4.5 and is not implemented; the paper notes there that two intuitive approaches to that case do not work.
Amsler, C., Prokhorov, A. and Schmidt, P. (2016). Endogeneity in stochastic frontier models. Journal of Econometrics, 190(2), 280--288.
Amsler, C., Prokhorov, A. and Schmidt, P. (2017). Endogenous environmental variables in stochastic frontier models. Journal of Econometrics, 199(2), 131--140.
Kutlu, L. (2010). Battese-Coelli estimator with endogenous regressors. Economics Letters, 109(2), 79--81.
endogeneity_test for the Wald test of whether the
correction was needed at all; sfm for the frontier without one.
# \donttest{
set.seed(1)
n <- 600
w1 <- rnorm(n); w2 <- rnorm(n); x1 <- rnorm(n)
eta <- rnorm(n)
v <- 0.5 * (0.6 * eta + sqrt(1 - 0.6^2) * rnorm(n))
x2 <- 0.9 * w1 - 0.7 * w2 + 0.5 * x1 + eta
y <- 0.5 + 0.8 * x1 - 0.6 * x2 + v - abs(rnorm(n))
dat <- data.frame(y = y, x1 = x1, x2 = x2, w1 = w1, w2 = w2)
fit <- ivsfm(y ~ x1 + x2, endogenous = ~x2, instruments = ~ w1 + w2,
data = dat)
fit$out
## Ignoring the endogeneity biases the coefficient on x2 toward zero.
sfm(y ~ x1 + x2, model_name = "NHN", data = dat)$out
# }
Run the code above in your browser using DataLab