Tests whether a lcsfm fit's latent classes are needed at all, against the null that one class describes the data. This is the first question a latent class fit raises and the ordinary likelihood ratio test cannot answer it, because under the null the second class's parameters are unidentified, the mixing proportion sits on a boundary, and the classes are interchangeable.
lcsfm_homogeneity(object, null = c("bootstrap", "chisq01"),
B = 199, c = 1, level = 0.05,
seed = NULL, quiet = FALSE, envir = parent.frame())An object of class c("sfa_mlrt", "htest"), so it prints like any other R test. Alongside statistic and p.value it carries null, B, boot (the bootstrap null draws), boot_ok, penalty_c, penalty_at_max, the penalised and unpenalised \(J\)-class log-likelihoods, logLik_1, class_prob, and reject.
An "sfareg" fit from lcsfm with model_name = "LCM" and n_class >= 2.
How to obtain the null distribution. "bootstrap" (the default) is a parametric bootstrap under the fitted one-class model. "chisq01" uses the published \(\chi^2_{0:1}\) limit and warns: see ‘Details’ for why it does not apply to "LCM" as implemented.
Bootstrap replications. Each is two refits, so this dominates the running time. The smallest attainable p-value is \(1/(B+1)\).
Tuning constant in the penalty \(c\log(J^J\prod_j p_j)\). Chen et al. (2001) report their simulations were insensitive to it; Stead, Wheat and Greene use 1 and 5.
Size used for the reported verdict.
Optional integer seed for the bootstrap. The RNG state is saved and restored, so a seeded call does not disturb the caller's stream.
Set TRUE to suppress the bootstrap progress messages.
Environment in which to re-evaluate the fit's data argument.
The statistic. lcsfm is refitted maximising the modified likelihood -- the ordinary log-likelihood plus Chen et al.'s penalty \(c\log(J^J\prod_j p_j)\), which diverges to \(-\infty\) as any class probability approaches 0 or 1 and so keeps every class occupied. The statistic is twice the difference between that maximum and the maximised log-likelihood of the one-class model, which is sfm(model_name = "NHN") on the same formula, since a single latent class is the ordinary frontier. The penalty is exactly zero at equal class probabilities, so the null model is recovered without a handicap.
Why the default null is a bootstrap. Stead, Wheat and Greene report the modified statistic is asymptotically \(\chi^2_{0:1}\), a 50:50 mixture of a point mass at zero and \(\chi^2_1\). That result is stated for the case in which a scalar parameter differs between classes -- Zhu and Zhang (2004) require \(\theta_1\) and \(\theta_2\) scalar, and the paper's own application holds every parameter except the noise scale common across classes.
lcsfm's "LCM" does not do that: it lets \(\sigma_v\), \(\sigma_u\) and the whole frontier vector vary by class. The limit therefore does not carry over, and this was measured rather than assumed. Over 200 replications of a one-class data generating process at \(n = 400\):
| nominal size | 10% | 5% |
actual rejection under "chisq01" | 82.0% | 63.5% |
with the sampling distribution tracking \(\chi^2_5\) -- exactly the parameter-count difference -- rather than \(\chi^2_{0:1}\), whose median is 0 against an observed 4.21.
A parametric bootstrap makes no appeal to an asymptotic distribution: it simulates from the fitted null and re-runs the whole procedure, so it is valid whatever the parameter structure. Over 60 replications of the same one-class process with B = 99:
| nominal size | 10% | 5% |
"chisq01" | 82.0% | 63.5% |
"bootstrap" | 10.0% | 6.7% |
both within a binomial standard error of nominal, with the p-values indistinguishable from uniform (KS against Uniform(0,1), \(p = 0.80\)). On a sample that "chisq01" rejected at \(p = 0.005\), the bootstrap returns \(p = 0.32\), and on genuinely two-class data the statistic is 17 times the largest of 99 null draws.
Cost. The bootstrap is \(B+1\) pairs of refits. Budget accordingly, and use seed for reproducibility.
"LCM_Z" is refused, by both this function and lcsfm(penalty_c = ): it makes the class probabilities depend on covariates, and the penalty is defined on a scalar mixing proportion.
Chen, H., Chen, J. and Kalbfleisch, J.D. (2001). A modified likelihood ratio test for homogeneity in finite mixture models. Journal of the Royal Statistical Society B, 63(1), 19--29.
Stead, A.D., Wheat, P. and Greene, W.H. (2023). On hypothesis testing in latent class and finite mixture stochastic frontier models, with application to a contaminated normal-half normal model. Journal of Productivity Analysis, 60(1), 37--48.
Zhu, H.-T. and Zhang, H. (2004). Hypothesis testing in mixture regression models. Journal of the Royal Statistical Society B, 66(1), 3--16.
lcsfm, TIC, skewness_test
# \donttest{
set.seed(21)
n <- 300
x1 <- rnorm(n)
cls <- rbinom(n, 1, 0.5)
d <- data.frame(y = ifelse(cls == 1, 4, 1) + x1 +
rnorm(n, 0, 0.4) - abs(rnorm(n, 0, 0.6)),
x1 = x1)
f <- lcsfm(y ~ x1, model_name = "LCM", data = d, n_class = 2)
## Small B for the example only; use the default in real work.
lcsfm_homogeneity(f, B = 19, seed = 1, quiet = TRUE)
# }
Run the code above in your browser using DataLab