Learn R Programming

sfa (version 1.2.0)

lcsfm_homogeneity: Test a latent class frontier against homogeneity

Description

Tests whether a lcsfm fit's latent classes are needed at all, against the null that one class describes the data. This is the first question a latent class fit raises and the ordinary likelihood ratio test cannot answer it, because under the null the second class's parameters are unidentified, the mixing proportion sits on a boundary, and the classes are interchangeable.

Usage

lcsfm_homogeneity(object, null = c("bootstrap", "chisq01"),
                  B = 199, c = 1, level = 0.05,
                  seed = NULL, quiet = FALSE, envir = parent.frame())

Value

An object of class c("sfa_mlrt", "htest"), so it prints like any other R test. Alongside statistic and p.value it carries null, B, boot (the bootstrap null draws), boot_ok, penalty_c, penalty_at_max, the penalised and unpenalised \(J\)-class log-likelihoods, logLik_1, class_prob, and reject.

Arguments

object

An "sfareg" fit from lcsfm with model_name = "LCM" and n_class >= 2.

null

How to obtain the null distribution. "bootstrap" (the default) is a parametric bootstrap under the fitted one-class model. "chisq01" uses the published \(\chi^2_{0:1}\) limit and warns: see ‘Details’ for why it does not apply to "LCM" as implemented.

B

Bootstrap replications. Each is two refits, so this dominates the running time. The smallest attainable p-value is \(1/(B+1)\).

c

Tuning constant in the penalty \(c\log(J^J\prod_j p_j)\). Chen et al. (2001) report their simulations were insensitive to it; Stead, Wheat and Greene use 1 and 5.

level

Size used for the reported verdict.

seed

Optional integer seed for the bootstrap. The RNG state is saved and restored, so a seeded call does not disturb the caller's stream.

quiet

Set TRUE to suppress the bootstrap progress messages.

envir

Environment in which to re-evaluate the fit's data argument.

Details

The statistic. lcsfm is refitted maximising the modified likelihood -- the ordinary log-likelihood plus Chen et al.'s penalty \(c\log(J^J\prod_j p_j)\), which diverges to \(-\infty\) as any class probability approaches 0 or 1 and so keeps every class occupied. The statistic is twice the difference between that maximum and the maximised log-likelihood of the one-class model, which is sfm(model_name = "NHN") on the same formula, since a single latent class is the ordinary frontier. The penalty is exactly zero at equal class probabilities, so the null model is recovered without a handicap.

Why the default null is a bootstrap. Stead, Wheat and Greene report the modified statistic is asymptotically \(\chi^2_{0:1}\), a 50:50 mixture of a point mass at zero and \(\chi^2_1\). That result is stated for the case in which a scalar parameter differs between classes -- Zhu and Zhang (2004) require \(\theta_1\) and \(\theta_2\) scalar, and the paper's own application holds every parameter except the noise scale common across classes.

lcsfm's "LCM" does not do that: it lets \(\sigma_v\), \(\sigma_u\) and the whole frontier vector vary by class. The limit therefore does not carry over, and this was measured rather than assumed. Over 200 replications of a one-class data generating process at \(n = 400\):

nominal size10%5%
actual rejection under "chisq01"82.0%63.5%

with the sampling distribution tracking \(\chi^2_5\) -- exactly the parameter-count difference -- rather than \(\chi^2_{0:1}\), whose median is 0 against an observed 4.21.

A parametric bootstrap makes no appeal to an asymptotic distribution: it simulates from the fitted null and re-runs the whole procedure, so it is valid whatever the parameter structure. Over 60 replications of the same one-class process with B = 99:

nominal size10%5%
"chisq01"82.0%63.5%
"bootstrap"10.0%6.7%

both within a binomial standard error of nominal, with the p-values indistinguishable from uniform (KS against Uniform(0,1), \(p = 0.80\)). On a sample that "chisq01" rejected at \(p = 0.005\), the bootstrap returns \(p = 0.32\), and on genuinely two-class data the statistic is 17 times the largest of 99 null draws.

Cost. The bootstrap is \(B+1\) pairs of refits. Budget accordingly, and use seed for reproducibility.

"LCM_Z" is refused, by both this function and lcsfm(penalty_c = ): it makes the class probabilities depend on covariates, and the penalty is defined on a scalar mixing proportion.

References

Chen, H., Chen, J. and Kalbfleisch, J.D. (2001). A modified likelihood ratio test for homogeneity in finite mixture models. Journal of the Royal Statistical Society B, 63(1), 19--29.

Stead, A.D., Wheat, P. and Greene, W.H. (2023). On hypothesis testing in latent class and finite mixture stochastic frontier models, with application to a contaminated normal-half normal model. Journal of Productivity Analysis, 60(1), 37--48.

Zhu, H.-T. and Zhang, H. (2004). Hypothesis testing in mixture regression models. Journal of the Royal Statistical Society B, 66(1), 3--16.

See Also

lcsfm, TIC, skewness_test

Examples

Run this code
# \donttest{
set.seed(21)
n   <- 300
x1  <- rnorm(n)
cls <- rbinom(n, 1, 0.5)
d   <- data.frame(y = ifelse(cls == 1, 4, 1) + x1 +
                      rnorm(n, 0, 0.4) - abs(rnorm(n, 0, 0.6)),
                  x1 = x1)

f <- lcsfm(y ~ x1, model_name = "LCM", data = d, n_class = 2)

## Small B for the example only; use the default in real work.
lcsfm_homogeneity(f, B = 19, seed = 1, quiet = TRUE)
# }

Run the code above in your browser using DataLab