Learn R Programming

sfa (version 1.2.0)

ivsfm: Stochastic Frontier Models with Endogenous Regressors

Description

Fits the stochastic frontier model of Amsler, Prokhorov and Schmidt (2016) in which one or more regressors is endogenous, in the simultaneous-equation sense: correlated with the statistical noise \(v\). Three estimators are provided -- full information maximum likelihood, a two-step control function, and corrected 2SLS. With uhet the model becomes that of Amsler, Prokhorov and Schmidt (2017), in which environmental variables entering the inefficiency scale may themselves be endogenous.

Usage

ivsfm(formula, endogenous, instruments, data,
      model_name = c("IVLIML", "IVCF", "C2SLS"), uhet = NULL,
      inefdec = TRUE, maxit.bobyqa = 10000, maxit.psoptim = 1000,
      maxit.optim = 1000, start_val = FALSE, PSopt = FALSE,
      optHessian = TRUE, Method = "L-BFGS-B", verbose = FALSE,
      rand.psoptim = NULL)

Value

An object of class "sfareg". out is a p x 3 matrix, one row per parameter, columns par, st_err and t-val. Parameters are the frontier coefficients, sigma_u, sigma_v, any delta_ terms from uhet, and one rho_ per endogenous variable (absent for "C2SLS"). Alongside the usual components:

Pi

Reduced-form coefficient matrix.

Sigma_xi

Reduced-form error covariance.

rho

Correlation of each reduced-form error with \(v\).

sigma_c

The conditional noise standard deviation \(\sigma_c\), which is smaller than sigma_v whenever \(\rho \ne 0\).

b_2sls

The 2SLS coefficients, which seed the likelihood and are the estimate "C2SLS" corrects.

wrong_skew

Whether the 2SLS residuals are skewed the wrong way.

jlms, efficiency

\(E[u_i \mid \varepsilon_i, \xi_i]\) and exp(-jlms).

vcov_rho

The full \(q \times q\) covariance of \(\hat\rho\), which endogeneity_test needs for the joint Wald statistic.

Standard errors for \(\rho\) arrived in 1.2.0. They were previously reported as NA, on the ground that \(\rho = t/\sqrt{1+t't}\) has a non-diagonal Jacobian and a delta-method value computed as though it were diagonal would be wrong. That much was true, but the Jacobian itself is short -- \(J = (I - \rho\rho')/s\) with \(s = \sqrt{1+t't}\) -- so \(\mathrm{Var}(\hat\rho) = J\,\mathrm{Var}(\hat t)\,J'\) is available exactly. Checked against a numerical Jacobian in the tests, and validated end-to-end by the size of endogeneity_test, which is 0.046 at a nominal 5% over 1000 replications.

Arguments

formula

The structural frontier, y ~ x1 + x2, listing all regressors including the endogenous ones.

endogenous

One-sided formula naming the endogenous variables, ~ x2. Each must appear in formula or in uhet.

instruments

One-sided formula of the excluded instruments, ~ w1 + w2. The included exogenous regressors are added automatically, so only the outside instruments are listed here. There must be at least as many as there are endogenous variables.

data

A data.frame holding every variable in the four formulas.

model_name

"IVLIML" (the default), "IVCF" or "C2SLS"; see ‘Details’.

uhet

Optional one-sided formula of environmental variables, giving \(\sigma_{U,i} = \sigma_U \exp(q_i'\delta)\). Supplying it selects the 2017 model; omitting it gives the 2016 one. Variables may appear in both uhet and endogenous.

inefdec

TRUE (default) for a production frontier, where inefficiency is subtracted; FALSE for a cost frontier.

maxit.bobyqa, maxit.psoptim, maxit.optim

Iteration caps for the three optimizer stages. Ignored by "C2SLS", which is closed-form.

start_val

Name the starting-value vector in the returned object.

PSopt

Run the particle-swarm stage between BOBYQA and optim.

optHessian

Compute the Hessian, and with it the standard errors.

Method

Method passed to optim for the final stage.

verbose

Report optimizer progress.

rand.psoptim

Optional seed for the particle-swarm stage.

Details

Unlike the rest of the package neither formula nor endogenous nor instruments takes a | segment. A pipe means “variance determinant” in sfm and psfm, and reusing it here for ivreg-style instrument syntax would give the same character two meanings across entry points, so a pipe is an error rather than being reinterpreted.

The model. Writing \(p_i\) for the endogenous block and \(z_i\) for the instruments, $$y_i = \beta'x_i + v_i - u_i, \qquad p_i = \Pi'z_i + \xi_i$$ $$u_i = \sigma_{U,i}|U_i|, \quad U_i \sim N(0,1), \qquad (v_i, \xi_i') \sim N(0, \Omega)$$ with \(u_i\) independent of \((v_i, \xi_i)\). Endogeneity is the correlation between \(\xi\) and \(v\): the regressors are correlated with the noise, not with the inefficiency.

Conditioning \(v\) on \(\xi\) turns this into an ordinary normal--half normal frontier in a shifted residual with a smaller noise variance, $$\mu_{c,i} = \Sigma_{v\xi}\Sigma_{\xi\xi}^{-1}\xi_i, \qquad \sigma_c^2 = \sigma_V^2 - \Sigma_{v\xi}\Sigma_{\xi\xi}^{-1}\Sigma_{\xi v}$$ which is what makes the likelihood closed-form and needs no simulation.

The estimators.

"IVLIML"

APS (2016) Equation (13): maximum likelihood over the frontier and the reduced form jointly, so \(\Pi\) and \(\Sigma_{\xi\xi}\) are estimated alongside \(\beta\). Efficient, and its standard errors are correct as reported.

"IVCF"

The two-step control function of Kutlu (2010), described in APS (2016) Section 4.4. \(\Pi\) and \(\Sigma_{\xi\xi}\) are fixed at their reduced-form least-squares values and only the frontier block is maximized. Cheaper, still consistent, but it discards the information about the reduced form carried in the frontier likelihood; its conventional standard errors understate the true ones unless \(\Sigma_{v\xi} = 0\). Measured over 30 replications at \(n = 1000\): with \(\rho = 0.6\) the mean reported standard error on the endogenous coefficient is 0.912 of the Monte Carlo standard deviation, while "IVLIML"'s is 1.005; with \(\rho = 0\) the two are identical, as the theory says they should be. Bootstrap, or use "IVLIML", when the standard errors matter.

"C2SLS"

APS (2016) Section 4.1: 2SLS, then \(\sigma_U\) and \(\sigma_V\) from the second and third moments of the residuals, then the intercept corrected by \(\sqrt{2/\pi}\,\hat\sigma_U\). The residuals use the actual endogenous regressors, not their fitted values, which the paper flags explicitly. Moment-based rather than likelihood-based, so the fit carries no $opt and logLik returns NA with a warning -- the same footing as psfm's "GTRE_SEQ1". It estimates no \(\rho\), since it never models the correlation it corrects for.

Inefficiency. jlms implements APS (2016) Equations (14)--(15), which condition on \(\xi\) as well as on \(\varepsilon\). Since \(\xi\) is correlated with \(v\) it carries information about \(u\) even though \(u\) is independent of \(\xi\), so this is a strictly better predictor than the plain Jondrow et al. (1982) formula.

What is assumed, and what is not covered. \(u\) is assumed independent of \((v, \xi)\). The case where the reduced-form error is correlated with the inefficiency as well requires the copula likelihood of APS (2016) Section 4.5 and is not implemented; the paper notes there that two intuitive approaches to that case do not work.

References

Amsler, C., Prokhorov, A. and Schmidt, P. (2016). Endogeneity in stochastic frontier models. Journal of Econometrics, 190(2), 280--288.

Amsler, C., Prokhorov, A. and Schmidt, P. (2017). Endogenous environmental variables in stochastic frontier models. Journal of Econometrics, 199(2), 131--140.

Kutlu, L. (2010). Battese-Coelli estimator with endogenous regressors. Economics Letters, 109(2), 79--81.

See Also

endogeneity_test for the Wald test of whether the correction was needed at all; sfm for the frontier without one.

Examples

Run this code
# \donttest{
set.seed(1)
n   <- 600
w1  <- rnorm(n); w2 <- rnorm(n); x1 <- rnorm(n)
eta <- rnorm(n)
v   <- 0.5 * (0.6 * eta + sqrt(1 - 0.6^2) * rnorm(n))
x2  <- 0.9 * w1 - 0.7 * w2 + 0.5 * x1 + eta
y   <- 0.5 + 0.8 * x1 - 0.6 * x2 + v - abs(rnorm(n))
dat <- data.frame(y = y, x1 = x1, x2 = x2, w1 = w1, w2 = w2)

fit <- ivsfm(y ~ x1 + x2, endogenous = ~x2, instruments = ~ w1 + w2,
             data = dat)
fit$out

## Ignoring the endogeneity biases the coefficient on x2 toward zero.
sfm(y ~ x1 + x2, model_name = "NHN", data = dat)$out
# }

Run the code above in your browser using DataLab