Learn R Programming

randomForestSRC (version 3.9.0)

multivariate.values: Extracting Multivariate Values

Description

Extract predictions, performance errors, variable importance (VIMP), and case-specific errors and VIMP from random forest fits and predictions. The functions combine response-specific results from multivariate regression, multivariate classification, and mixed-outcome forests. They also work with univariate forests. get.mv.formula() constructs a multivariate formula from response names.

Usage

get.mv.predicted(obj, oob = TRUE)

get.mv.error(obj, standardize = FALSE, pretty = TRUE, block = FALSE)

get.mv.error.block(obj, standardize = FALSE)

get.mv.vimp(obj, standardize = FALSE, pretty = TRUE)

get.mv.cserror(obj, standardize = FALSE)

get.mv.csvimp(obj, standardize = FALSE)

get.mv.formula(ynames)

Value

get.mv.predicted

A numeric matrix with one row per observation and columns for response predictions, class probabilities, or event-specific predictions. A single-column result is also a matrix.

get.mv.error

A named numeric vector by default, or a list named by response when pretty = FALSE. Entries contain the final error or error row. With block = TRUE, the result is a list of block-error sequences. Returns NULL when no response has an error result.

get.mv.error.block

A list of block-error vectors or matrices named by response, or NULL when no response has block errors.

get.mv.vimp

A numeric matrix by default, or a list of matrices named by response when pretty = FALSE. Returns NULL if the first response has no importance result.

get.mv.cserror

Case-specific error values for one response, with their original vector or array dimensions, or a list named by response for multiple responses. Returns NULL for survival families or if the first response has no case-specific error result.

get.mv.csvimp

A case-by-variable matrix for one response, or a list of these matrices named by response for multiple responses. Returns NULL for survival families or if the first response has no case-specific VIMP result.

get.mv.formula

An R formula with the supplied response names on the left of ~ and a dot on the right.

Arguments

obj

An object returned by rfsrc() or predict.rfsrc(). For multivariate and mixed-outcome objects, obj$yvar must contain the response columns with their original types and factor levels.

oob

If TRUE, use predicted.oob for each response, falling back to predicted when that response's OOB component is NULL. If FALSE, use predicted.

standardize

If TRUE, divide regression errors and VIMP by the corresponding response variance in obj$yvar. Classification and survival values are unchanged. See Standardization for the treatment of zero or missing variances.

pretty

If TRUE, return errors as a named vector and VIMP as a matrix, using only the overall all value for each classification response. If FALSE, return a list named by response, including class-specific values. Competing-risk event-specific values are included in both formats. Values are not rounded.

block

If TRUE, return the block-error sequence from err.block.rate instead of the final error from err.rate. This sets pretty = FALSE. get.mv.error.block() is shorthand for this choice.

ynames

Character vector of response column names for get.mv.formula(). These columns must be in the data used to fit the forest.

Details

The extraction functions use results already computed during fitting or prediction. Results are reported separately for each response. Subset the returned vector, matrix, or list to select responses or predictors.

Predictions

get.mv.predicted() returns a matrix with observations in rows and predictions in columns:

  • Regression: one column per response.

  • Classification: one probability column per class, named response.class.

  • Right-censored survival: the mortality prediction.

  • Competing risks: one prediction column per event, named response.event.

Responses follow obj$yvar.names; class and event columns follow their order in the object. Time-indexed survival, cumulative hazard, and cumulative incidence arrays are not included.

By default, OOB predictions are used for each response when its predicted.oob component is present. Missing values within that component remain missing, even if every value is missing. The function uses predicted only when the entire OOB component is NULL. Set oob = FALSE to use full-ensemble or new-data predictions.

Performance errors and variable importance

get.mv.error() returns the final performance error for each response, using the error measure selected during fitting or prediction. With pretty = FALSE, classification results include all and the class-specific errors. Survival results include all error columns, including event-specific errors for competing risks.

get.mv.error.block() returns the full block-error sequence computed during fitting or prediction. It is equivalent to get.mv.error(obj, standardize = standardize, block = TRUE); the block size is determined when those errors are computed.

get.mv.vimp() extracts variable importance. Request importance when fitting the forest or use vimp before calling this function. With pretty = TRUE, the result is a matrix with predictors in rows and response results in columns. With pretty = FALSE, each response has its own matrix, including classification all and class-specific columns and competing-risk event columns.

Case-specific values

get.mv.cserror() calculates case-specific error as cse.num / cse.den; get.mv.csvimp() calculates case-specific VIMP as csv.num / csv.den. These numerator and denominator components must be saved in the fitted or prediction object. Case-specific error is calculated from these components, not by applying a loss to predicted.oob and the observed response.

With one response, values are returned directly; with multiple responses, they are returned in a list named by response. Case-specific VIMP has observations in rows and VIMP variables in columns, with variable names taken from importance when available. Both functions return NULL for right-censored survival and competing risks.

Standardization

For a regression response \(Y\), standardize = TRUE divides each error or VIMP value by var(Y, na.rm = TRUE). The variance uses the response values in obj$yvar: training responses for a fitted object and evaluation responses for a test-prediction object. Classification and survival values are unchanged. Standardization is off by default and does not apply to predictions.

get.mv.error(), get.mv.error.block(), and get.mv.vimp() divide by the variance directly, so a zero or missing variance can produce nonfinite values. The case-specific functions use a divisor of one when the variance is zero or NA.

Missing results

A NULL component means that the corresponding result is absent from obj. An NA entry is a missing value within an existing result.

  • get.mv.error() returns NULL when no response has an error result. Otherwise, a response without an error is represented by NA in vector output or NULL in list output. The same list behavior applies to get.mv.error.block().

  • get.mv.vimp(), get.mv.cserror(), and get.mv.csvimp() return NULL if the first response lacks the corresponding result, even when later responses have results. If the first response has a result, missing later results are NULL in list output.

Constructing a multivariate formula

get.mv.formula(ynames) returns Multivar(y1, y2, ...) ~ .. Responses may be continuous, factors, or a mixture; their types are determined from the data when the forest is fitted. The dot specifies the remaining data columns as predictors.

See Also

rfsrc, predict.rfsrc, vimp, subsample, classification.performance

Examples

Run this code
## ------------------------------------------------------------
## A basic multivariate analysis
## ------------------------------------------------------------
o <- rfsrc(cbind(Ozone, Temp) ~ ., data = na.omit(airquality))
print(head(get.mv.predicted(o)))
print(get.mv.error(o))

# \donttest{
## ------------------------------------------------------------
## Select a response from the stored results
## ------------------------------------------------------------
pred.oob <- get.mv.predicted(o)
print(head(pred.oob[, "Temp", drop = FALSE]))
print(get.mv.error(o)["Temp"])
print(get.mv.error(o, standardize = TRUE))
print(head(get.mv.predicted(o, oob = FALSE)))

## ------------------------------------------------------------
## Formula construction, VIMP, and block errors
## ------------------------------------------------------------
f <- get.mv.formula(c("Ozone", "Temp"))
print(f)

o.vimp <- rfsrc(f, data = na.omit(airquality), ntree = 100,
                importance = "permute", block.size = 10)
print(get.mv.vimp(o.vimp))
print(get.mv.vimp(o.vimp, standardize = TRUE))
print(get.mv.vimp(o.vimp, pretty = FALSE)[["Temp"]])
print(head(get.mv.error.block(o.vimp)[["Temp"]]))

## Optional case-specific values are NULL when not saved in the object.
print(get.mv.cserror(o.vimp))
print(get.mv.csvimp(o.vimp))

## ------------------------------------------------------------
## Mixed outcomes: include class-specific entries
## ------------------------------------------------------------
f.mix <- get.mv.formula(c("Sepal.Length", "Species"))
mix <- rfsrc(f.mix, data = iris, ntree = 100,
             importance = "permute", block.size = 10)
print(colnames(get.mv.predicted(mix)))
print(get.mv.error(mix))
print(get.mv.error(mix, pretty = FALSE)[["Species"]])
print(get.mv.vimp(mix, pretty = FALSE)[["Species"]])

## ------------------------------------------------------------
## Extract predictions and errors for held-out observations
## ------------------------------------------------------------
dta <- na.omit(airquality)
set.seed(17)
train <- sample(seq_len(nrow(dta)), floor(.7 * nrow(dta)))
fit <- rfsrc(f, data = dta[train, ], ntree = 100)
p.test <- predict(fit, newdata = dta[-train, ])
print(head(get.mv.predicted(p.test, oob = FALSE)))
print(get.mv.error(p.test))
# }

Run the code above in your browser using DataLab