Learn R Programming

sfa (version 1.2.0)

gof_test: Goodness-of-fit test for the assumed inefficiency distribution

Description

Tests the distributional assumption on \(u\), holding the normality of \(v\) fixed, following Wang, Amsler and Schmidt (2011).

Usage

gof_test(object, data = NULL, test = c("ks", "chisq"),
         null = c("bootstrap", "asymptotic"), B = 199, cells = 10,
         seed = NULL)

Value

A data frame with one row per test: test, statistic, null, p.value and B (replications actually used).

Arguments

object

An "sfareg" fit from sfm.

data

The data the model was fitted to. sfm() does not retain it, so it must be supplied or be findable from the fit's call.

test

"ks", "chisq", or both.

null

"bootstrap" for the parametric bootstrap of Wang, Amsler and Schmidt, or "asymptotic".

B

Bootstrap replications. Each one refits the model.

cells

Number of equiprobable cells for the chi-square test.

seed

Optional seed for the bootstrap.

Details

The distribution of \(u\) is the assumption in this model least often defended and least often tested, and it is testable: with the noise distribution maintained, the assumed \(u\) implies a distribution for the composed error, so rejecting that is rejecting the assumed \(u\).

The test is on \(\varepsilon\), not on \(\hat u\). Wang, Amsler and Schmidt point out that \(\hat u = E[u \mid \varepsilon]\) is a monotonic function of \(\varepsilon\), so the KS test is identical either way and the chi-square test is identical when the cells are defined conformably -- but \(\varepsilon\) is much easier to work with. It is a mistake, not a diagnostic, to compare the observed spread of \(\hat u\) with the assumed density of \(u\): those are different distributions.

Parameter estimation cannot be ignored. Both statistics are evaluated at \(\hat\theta\), which changes their null distributions. The default "bootstrap" handles this by copying the estimation step exactly: each replication draws a composed error from the fitted model, rebuilds the response, refits, forms its own residuals, and recomputes the statistic. "asymptotic" returns a \(\chi^2(k-1-m)\) p-value for the chi-square statistic, which the authors note is conservative when evaluated at the MLE -- that reference distribution belongs to the minimum-chi-square estimator. There is no asymptotic KS p-value here: with \(\theta\) estimated the Kolmogorov distribution does not apply, and Bai's (2003) martingale transformation is not implemented, so NA is returned rather than a number that would look valid.

Size and power. Measured in this package over 200 replications with B = 49, half-normal data fitted as half-normal for size and exponential data fitted as half-normal for power:

\(n\)KS .10KS .05chi2 .10chi2 .05
size2000.0650.0250.0850.065
size8000.0700.0400.0650.010
power2000.5050.3950.2350.145
power8000.9850.9650.8100.655
power20001.0001.0000.9900.985

Both tests hold their size, mildly conservatively. The KS test dominates the chi-square test at every sample size, which is the paper's own conclusion and the reason it is the better default: the chi-square statistic discards the within-cell information that KS uses. Power against the half-normal/exponential pair -- one of the harder distinctions, since both are one-parameter and similarly shaped -- is about 0.4 at \(n = 200\) and essentially 1 by \(n = 800\), so a failure to reject in a small sample is weak evidence.

References

Wang, W. S., Amsler, C. and Schmidt, P. (2011). Goodness of fit tests in stochastic frontier models. Journal of Productivity Analysis 35, 95--118.

Bai, J. (2003). Testing parametric conditional distributions of dynamic models. Review of Economics and Statistics 85, 531--549.

See Also

spec_test, pcomposed_model, sfm

Examples

Run this code
d <- as.data.frame(data_gen_cs(N = 200, rand = 5, sig_u = 1, sig_v = 0.5,
                               cons = 0.5, beta1 = 0.5, beta2 = 0.5,
                               a = 5, mu = 0.1))
f <- sfm(y_pcs ~ x1 + x2, model_name = "NHN", data = d)
# \donttest{
## Every bootstrap replication REFITS, so keep B small in an example; use
## B = 999 for a real test.
gof_test(f, data = d, B = 39, seed = 1)
# }
gof_test(f, data = d, test = "chisq", null = "asymptotic")

Run the code above in your browser using DataLab