Learn R Programming

randomForestSRC (version 3.9.0)

classification.performance: Classification Performance Metrics

Description

Evaluate predicted classes and probabilities using confusion matrices, class-specific misclassification rates, ROC area under the curve (AUC), Brier score, log loss, and binary precision-recall summaries. get.bayes.rule() assigns classes from predicted probabilities.

Usage

get.confusion(y, class.or.prob)

get.misclass.error(y, yhat)

get.auc(y, prob)

get.brier.error(y, prob, normalized = TRUE, vector = FALSE)

get.logloss(y, prob, robust = TRUE)

get.pr.auc(truth, yhat)

get.pr.curve(truth, yhat)

get.bayes.rule(prob, class.relfrq = NULL)

Value

get.confusion

A numeric matrix of class counts with an additional class.error column.

get.misclass.error

An unnamed numeric vector of class-specific error rates, in sort(unique(y)) order.

get.auc

A scalar AUC, or NA when unavailable.

get.brier.error

A scalar Brier score by default, or an unnamed length-\(J\) vector of scaled class contributions in probability-column order when vector = TRUE.

get.logloss

A scalar mean log loss. It can be infinite with robust = FALSE, or NaN when no losses remain for averaging.

get.pr.auc

An unnamed numeric vector of length two: the model area followed by the random reference area. Both entries are NA when the calculation is unavailable.

get.pr.curve

A numeric matrix with columns recall, precision, and threshold, or NULL when the calculation is unavailable.

get.bayes.rule

A factor of predicted classes with levels given by the probability-column names.

Arguments

y

Observed class labels, usually a factor with levels in probability-column order. get.auc() also accepts nonfactor labels, with columns in sort(unique(y)) order. get.misclass.error() also accepts vectors of class labels.

prob

Numeric matrix with one row per observation and one column per class, in response-level order. Name columns with the class labels. get.auc() uses column positions; get.logloss() and get.bayes.rule() require column names. get.confusion() and get.brier.error() supply missing names from levels(y).

class.or.prob

A factor of predicted labels with the same levels as y, or a probability matrix as described for prob.

yhat

For get.misclass.error(), predicted class labels aligned with y. For precision-recall summaries, a numeric score vector or a matrix or data frame with one or two columns. A vector or single column gives positive-class scores; larger values favor that class. Scores need not lie in \([0,1]\). See Details for two-column inputs.

truth

Binary responses coded 0/1, or a two-level factor. For 0/1 labels, 1 is positive. For other factors, the less frequent class is positive, with ties resolved by the first factor level. The positive class is selected before excluding missing scores.

normalized

If TRUE, scale the Brier score so that equal class probabilities give score one. If FALSE, average the squared probability errors over classes. See Details.

vector

If TRUE, return scaled Brier contributions for each class instead of their sum.

robust

For get.logloss(), exclude infinite losses before averaging. Probabilities are unchanged.

class.relfrq

Optional two-class relative frequencies, in probability-column order. NULL selects the largest-probability rule; supplying frequencies selects the minority-prevalence rule described below.

Details

Supplying predictions

Supply responses and predictions in matching row order. For OOB evaluation of a grow object o, use o$yvar with o$predicted.oob for probabilities or o$class.oob for predicted classes. For new data, use the prediction object's yvar with predicted or class.

To compare metrics on the same observations, retain observed labels and finite probabilities. get.logloss() and get.misclass.error() require nonmissing observed labels.

For a combined binary summary with sensitivity, specificity, F1, G-mean, and random-reference comparisons, use get.imbalanced.performance.

Class assignments and misclassification

With class.relfrq = NULL, get.bayes.rule() selects the class with the largest probability, breaking ties at random. Rows with all probabilities missing receive NA. With two class frequencies supplied, it assigns the minority class when its probability is at least its relative frequency, and the majority class otherwise. If frequencies tie, the class in the first probability column is treated as the minority. Supply finite probabilities for this rule.

get.confusion() tabulates observed classes in rows and predicted classes in columns. The final class.error column gives each observed class's misclassification rate, rounded to four decimal places. Missing response/prediction pairs are excluded. Probability matrices use the largest-probability rule. For RFQ or another thresholded rule, supply the predicted factor instead, such as o$class.oob.

get.misclass.error() returns one error rate per observed class, in sort(unique(y)) order. A missing predicted label makes the corresponding class rate unavailable.

ROC area under the curve

get.auc() computes the multiclass AUC of Hand and Till (2001). For each class pair, it averages two rank-based AUCs, one using each class's probability as the score, then averages across available pairs. For complementary binary probabilities, the two directions agree. Larger values indicate better discrimination.

Missing labels and nonfinite scores are excluded from each rank calculation. Each direction requires at least two finite scores per class. Unavailable pairs are omitted; the result is NA when none remain.

Brier score

The Brier score measures squared probability error (Brier, 1950). For \(J\) classes, let $$b_j=\mathrm{mean}_i\left[(I(y_i=j)-p_{ij})^2\right].$$ get.brier.error() returns $$\frac{J}{J-1}\sum_{j=1}^{J} b_j$$ for normalized = TRUE, or $$\frac{1}{J}\sum_{j=1}^{J} b_j$$ for normalized = FALSE. Smaller values are better. Equal probabilities \(p_{ij}=1/J\) give normalized score one. For complementary binary probabilities, the unnormalized score is the mean squared error of either class probability; the normalized score is four times that value.

With vector = TRUE, each element is \(b_j\) multiplied by the selected scaling constant. Missing losses are omitted within each class. The scalar sums available contributions, returning NA when none remain.

Log loss

With every response level represented, get.logloss() averages the negative log probability of the observed class, \(-\log(p_{i,y_i})\), using the natural logarithm (see Gneiting and Raftery, 2007). Smaller values are better. Missing losses are omitted. With robust = TRUE, infinite losses from zero probabilities are also omitted; robust = FALSE retains them. Probabilities are not clipped. An unused factor level contributes one zero term to the average.

Precision-recall summaries

For two-column scores, named columns match the original response levels; unnamed columns follow factor-level order, or 0, 1 order for nonfactor responses. The positive-class column is selected. Rows with a missing response or nonfinite selected score are excluded; both classes must remain.

get.pr.auc() returns the precision-recall area, calculated using analytic precision-recall interpolation, and a random reference area equal to the positive-class proportion among scored observations. get.pr.curve() returns recall, precision, and thresholds in decreasing recall order. Larger precision-recall areas are better. See Davis and Goadrich (2006) for background on precision-recall curves.

References

Brier, G.W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78, 1-3. tools:::Rd_expr_doi("10.1175/1520-0493(1950)078<0001:vofeit>2.0.CO;2")

Davis, J. and Goadrich, M. (2006). The relationship between precision-recall and ROC curves. Proceedings of the 23rd International Conference on Machine Learning, 233-240. tools:::Rd_expr_doi("10.1145/1143844.1143874")

Gneiting, T. and Raftery, A.E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102, 359-378. tools:::Rd_expr_doi("10.1198/016214506000001437")

Hand, D.J. and Till, R.J. (2001). A simple generalisation of the area under the ROC curve for multiple class classification problems. Machine Learning, 45, 171-186. tools:::Rd_expr_doi("10.1023/A:1010920819831")

See Also

rfsrc, predict.rfsrc, imbalanced.rfsrc, get.imbalanced.performance, get.brier.survival, get.auct.survival

Examples

Run this code
## ------------------------------------------------------------
## A basic calculation from observed labels and probabilities
## ------------------------------------------------------------
y <- factor(c("no", "no", "no", "no", "yes", "yes"),
            levels = c("no", "yes"))
p <- c(.10, .20, .65, .35, .40, .85)
prob <- cbind(no = 1 - p, yes = p)
print(get.confusion(y, prob))
print(get.auc(y, prob))
print(get.brier.error(y, prob))

# \donttest{
## ------------------------------------------------------------
## Class-specific errors and probability losses
## ------------------------------------------------------------
yhat <- get.bayes.rule(prob)
print(setNames(get.misclass.error(y, yhat), levels(y)))
print(get.brier.error(y, prob, normalized = FALSE))
print(setNames(get.brier.error(y, prob, vector = TRUE), colnames(prob)))
print(get.logloss(y, prob))

## Use supplied class frequencies for the binary RFQ decision rule.
class.frq <- as.numeric(prop.table(table(y)))
yhat.rfq <- get.bayes.rule(prob, class.relfrq = class.frq)
print(get.confusion(y, yhat.rfq))

## ------------------------------------------------------------
## Precision-recall from a score vector
## ------------------------------------------------------------
truth <- as.integer(y == "yes")
pr.auc <- get.pr.auc(truth, p)
print(setNames(pr.auc, c("model", "random")))
pr <- get.pr.curve(truth, p)
plot(pr[, "recall"], pr[, "precision"], type = "l",
     xlim = c(0, 1), ylim = c(0, 1),
     xlab = "Recall", ylab = "Precision")
abline(h = pr.auc[2], lty = 2)

## A single score column is also accepted.
print(get.pr.auc(truth, matrix(p, ncol = 1)))

## ------------------------------------------------------------
## Multiclass forest: OOB and test-data performance
## ------------------------------------------------------------
set.seed(17)
train <- c(1:35, 51:85, 101:135)
o <- rfsrc(Species ~ ., data = iris[train, ], ntree = 100)

## Select one common set of observed responses and finite OOB scores.
p.oob <- o$predicted.oob
keep <- !is.na(o$yvar) & rowSums(!is.finite(p.oob)) == 0
print(get.confusion(o$yvar[keep], o$class.oob[keep]))
print(get.auc(o$yvar[keep], p.oob[keep, , drop = FALSE]))

## Use the current responses and probabilities for test-data scoring.
p.test <- predict(o, newdata = iris[-train, ])
print(c(
  auc = get.auc(p.test$yvar, p.test$predicted),
  brier = get.brier.error(p.test$yvar, p.test$predicted),
  logloss = get.logloss(p.test$yvar, p.test$predicted)
))
# }

Run the code above in your browser using DataLab