The distribution of \(u\) is the assumption in this model least often
defended and least often tested, and it is testable: with the noise
distribution maintained, the assumed \(u\) implies a distribution for the
composed error, so rejecting that is rejecting the assumed \(u\).
The test is on \(\varepsilon\), not on \(\hat u\). Wang, Amsler
and Schmidt point out that \(\hat u = E[u \mid \varepsilon]\) is a
monotonic function of \(\varepsilon\), so the KS test is identical either
way and the chi-square test is identical when the cells are defined
conformably -- but \(\varepsilon\) is much easier to work with. It is a
mistake, not a diagnostic, to compare the observed spread of \(\hat u\)
with the assumed density of \(u\): those are different distributions.
Parameter estimation cannot be ignored. Both statistics are evaluated
at \(\hat\theta\), which changes their null distributions. The default
"bootstrap" handles this by copying the estimation step exactly: each
replication draws a composed error from the fitted model, rebuilds the
response, refits, forms its own residuals, and recomputes the
statistic. "asymptotic" returns a \(\chi^2(k-1-m)\) p-value for the
chi-square statistic, which the authors note is conservative when
evaluated at the MLE -- that reference distribution belongs to the
minimum-chi-square estimator. There is no asymptotic KS p-value here: with
\(\theta\) estimated the Kolmogorov distribution does not apply, and Bai's
(2003) martingale transformation is not implemented, so NA is returned
rather than a number that would look valid.
Size and power. Measured in this package over 200 replications with
B = 49, half-normal data fitted as half-normal for size and
exponential data fitted as half-normal for power:
| \(n\) | KS .10 | KS .05 | chi2 .10 | chi2 .05 |
| size | 200 | 0.065 | 0.025 | 0.085 | 0.065 |
| size | 800 | 0.070 | 0.040 | 0.065 | 0.010 |
| power | 200 | 0.505 | 0.395 | 0.235 | 0.145 |
| power | 800 | 0.985 | 0.965 | 0.810 | 0.655 |
| power | 2000 | 1.000 | 1.000 | 0.990 | 0.985 |
Both tests hold their size, mildly conservatively. The KS test
dominates the chi-square test at every sample size, which is the paper's own
conclusion and the reason it is the better default: the chi-square statistic
discards the within-cell information that KS uses. Power against the
half-normal/exponential pair -- one of the harder distinctions, since both
are one-parameter and similarly shaped -- is about 0.4 at \(n = 200\) and
essentially 1 by \(n = 800\), so a failure to reject in a small sample is
weak evidence.