Implementation of the Hull method suggested by Lorenzo-Seva, Timmerman, and Kiers (2011), with an extension to principal axis factoring. See details for parallelization.
HULL(
x,
N = NA,
n_fac_theor = NA,
method = c("PAF", "ULS", "ML"),
gof = c("CAF", "CFI", "RMSEA"),
eigen_type = c("SMC", "PCA", "EFA"),
use = c("pairwise.complete.obs", "all.obs", "complete.obs", "everything",
"na.or.complete"),
cor_method = c("pearson", "spearman", "kendall", "poly", "tetra"),
n_datasets = 1000,
percent = 95,
decision_rule = c("means", "percentile", "crawford"),
n_factors = 1,
...
)An object of class efa_retention (see print.efa_retention() and
plot.efa_retention() for the print and plot methods). Its main fields are:
A named numeric vector with the suggested number of factors
for each requested goodness-of-fit index ("CAF", "CFI", and/or
"RMSEA").
A list with one record per goodness-of-fit index, each holding the goodness-of-fit values, the degrees of freedom, the hull membership, and the retained solution used for printing and plotting.
A list of the settings used, including n_fac_max, the upper
bound J of the number of factors to extract (see details).
matrix or data.frame. Dataframe or matrix of raw data or matrix with correlations.
numeric. Number of cases in the data. This is passed to PARALLEL. Only has to be specified if x is a correlation matrix, otherwise it is determined based on the dimensions of x.
numeric. Theoretical number of factors to retain. One plus the larger of this number and the number of factors suggested by PARALLEL is used as the upper bound J of factors to extract in the Hull method.
character. The estimation method to use. One of "PAF",
"ULS", or "ML", for principal axis factoring, unweighted
least squares, and maximum likelihood, respectively.
character. The goodness of fit index to use. Either "CAF",
"CFI", or "RMSEA", or any combination of them.
If method = "PAF" is used, only
the CAF can be used as goodness of fit index. For details on the CAF, see
Lorenzo-Seva, Timmerman, and Kiers (2011).
character. On what the eigenvalues should be found in the
parallel analysis. Can be one of "SMC", "PCA", or "EFA".
If using "SMC" (default), the diagonal of the correlation matrices is
replaced by the squared multiple correlations (SMCs) of the indicators. If
using "PCA", the diagonal values of the correlation
matrices are left to be 1. If using "EFA", eigenvalues are found on the
correlation matrices with the final communalities of an EFA solution as
diagonal. This is passed to PARALLEL().
character. Passed to stats::cor() if raw data
is given as input. Default is "pairwise.complete.obs".
character. One of "pearson", "spearman", or "kendall",
passed to stats::cor(). "poly" and "tetra" are not supported because
HULL derives its factor-search bound from an internal parallel analysis
against continuous reference data.
Default is "pearson".
numeric. The number of datasets to simulate. Default is 1000.
This is passed to PARALLEL().
numeric. The percentile to take from the simulated eigenvalues.
Default is 95. This is passed to PARALLEL().
character. Which rule to use to determine the number of
factors to retain. Default is "means", which will use the average
simulated eigenvalues. "percentile", uses the percentiles specified
in percent. "crawford" uses the 95th percentile for the first factor
and the mean afterwards (based on Crawford et al, 2010). This is passed to PARALLEL().
numeric. Number of factors to extract if "EFA" is
included in eigen_type. Default is 1. This is passed to
PARALLEL().
Further arguments passed to EFA(), also in
PARALLEL().
The Hull method aims to find a model with an optimal balance between
model fit and number of parameters, retaining only major factors
(Lorenzo-Seva, Timmerman, & Kiers, 2011). It fits 0 to J factors -- where
J is the number of factors suggested by parallel analysis (or n_fac_theor,
if that is larger), plus one -- keeps the solutions on the upper boundary of
the convex hull of goodness-of-fit against degrees of freedom, and selects the
one at the sharpest elbow, i.e. with the highest st value.
The PARALLEL function and the principal axis factoring of the
different number of factors can be parallelized using the future framework,
by calling the future::plan() function. The examples
provide example code on how to enable parallel processing.
Note that if gof = "RMSEA" is used, 1 - RMSEA is actually used to
compare the different solutions. Thus, the threshold of .05 is then .95. This
is necessary due to how the heuristic to locate the elbow of the hull works.
The ML estimation method uses the psych::fa()
starting values. See also the EFA documentation.
N_FACTORS() as a wrapper function for this and the other factor
retention criteria.
Other factor retention criteria:
CD(),
EKC(),
KGC(),
MAP(),
NEST(),
PARALLEL(),
SCREE(),
SMT()
# \donttest{
# using PAF (this will throw a warning if gof is not specified manually
# and CAF will be used automatically)
HULL(test_models$baseline$cormat, N = 500, gof = "CAF")
# using ML with all available fit indices (CAF, CFI, and RMSEA)
HULL(test_models$baseline$cormat, N = 500, method = "ML")
# using ULS with only RMSEA
HULL(test_models$baseline$cormat, N = 500, method = "ULS", gof = "RMSEA")
# }
if (FALSE) {
# using parallel processing (Note: plans can be adapted, see the future
# package for details)
future::plan(future::multisession)
HULL(test_models$baseline$cormat, N = 500, gof = "CAF")
}
Run the code above in your browser using DataLab