Learn R Programming

randomForestSRC (version 3.9.0)

impute.learn.rfsrc: Learn a predictive imputer for test-time imputation and OOD scoring

Description

Learns a predictive imputer from training data for later use on new data.

If the training data contain missing values, the function first imputes them using impute. It then fits one saved full-sweep learner per selected target on the completed training data and reuses those learners later to update missing values in new data without refitting on the test set.

The same saved learner bank can also be used to score new data for out-of-distribution (OOD) behavior. Note that OOD scores are available even when new data have missing values. Each selected target is reconstructed from its saved conditional learner and compared with the observed value. Target-wise discrepancies are calibrated against a training reference calculated from out-of-bag predictions computed during training.

If the training data are complete and target.mode = "all", the initial training-data imputation step is skipped and the full-sweep learners are fit directly from the complete training data.

If supervised.formula is supplied, the function also fits an internal supervised forest from the training data. The supervised forest is fit after training imputation and provides auxiliary predictors for test-time imputation and OOD scoring. The auxiliary predictors use supervised information learned from the training outcomes. Supervised outcomes supplied with new data are dropped and are not used at deployment time. Leave supervised.formula unspecified to use the learned imputer without these auxiliary predictors.

Usage

impute.learn.rfsrc(formula, data,
  ntree = 100, nodesize = 1, nsplit = 10,
  nimpute = 2, fast = FALSE, blocks,
  mf.q, max.iter = 10, eps = 0.01,
  ytry = NULL, always.use = NULL, verbose = TRUE,
  ...,
  supervised.formula = NULL,
  supervised.args = list(),
  full.sweep.options = list(ntree = 100, nsplit = 10),
  target.mode = c("missing.only", "all"),
  deployment.xvars = NULL,
  anonymous = TRUE,
  learner.prefix = "impute.learner.",
  learner.root = "learners",
  out.dir = NULL,
  wipe = TRUE,
  keep.models = is.null(out.dir),
  keep.ximp = FALSE,
  save.on.fit = !is.null(out.dir),
  save.ood = TRUE,
  weight = NULL)

save.impute.learn.rfsrc(object, path, wipe = TRUE, verbose = TRUE)

load.impute.learn.rfsrc(path, targets = NULL, lazy = TRUE, verbose = TRUE)

# S3 method for impute.learn.rfsrc predict(object, newdata, max.predict.iter = 3L, eps = 1e-3, targets = NULL, restore.integer = TRUE, cache.learners = c("session", "none", "all"), verbose = TRUE, ...)

impute.ood.rfsrc(object, newdata, targets = NULL, max.predict.iter = 3L, eps = 1e-3, cache.learners = c("all", "session", "none"), weight = NULL, aggregate = c("bounded.product", "weighted.mean", "weighted.lp", "weighted.lp.log", "top.k"), aggregate.args = list(), return.details = FALSE, return.reconstruction = FALSE, verbose = TRUE, ...)

Value

impute.learn returns an object of class

c("impute.learn.rfsrc", "impute.learn"). The object contains a manifest, optionally the fitted full-sweep learners, optionally the internal supervised forest when

supervised.formula is used, optionally the completed raw training predictor table, and optionally a path to the saved imputer on disk. If save.ood = TRUE, the manifest also contains an

ood component storing compact target-wise OOD references, the saved row-by-target training OOD score matrix used for later percentile recalibration, and the default OOD aggregation weights. When supervised mode is active, the manifest also records the supervised family, response names, and the internally created auxiliary predicted.* variable names.

load.impute.learn returns an object of the same class.

predict.impute.learn returns a data frame with imputed values overlaid on the raw predictor table, retaining its row names. An attribute named

"impute.learn.info" contains prediction-time diagnostics such as the number of sweep passes, pass-difference history, caching mode, disk-load counts, schema harmonization details, dropped supervised response columns when present, row-wise unseen-factor flags, supervised-auxiliary diagnostics when present, and any targets skipped because a learner was unavailable or a prediction failed.

pass.updated.cells and pass.failed.cells count accepted updates and failed updates in each pass; converged and

stopping.reason distinguish convergence from initialization-only, empty-input, iteration-limit, and failed-update stopping.

conversion.issues records the row indices of nonmissing values that became missing during numeric conversion. In both prediction and OOD diagnostics, n.disk.loads counts successful target-learner load operations, including repeated loads with cache.learners = "none".

disk.load.targets lists the distinct targets loaded. Supervised forest loading is reported separately in info$supervised and is not included in the target-learner count.

impute.ood returns an object of class

c("impute.ood.rfsrc", "impute.ood"). It is a list with the following components:

  • score: the row-level aggregate of calibrated target-wise OOD scores under the requested aggregate and weight. Larger values indicate greater out-of-distribution behavior. For aggregate = "weighted.lp.log", the probability floor bounds the mathematical aggregate by \(-\log(\varepsilon)\) for \(0 < \varepsilon < 1\). With the default eps = 1e-12, the upper bound is approximately 27.63.

  • score.percentile: the percentile of score relative to a row-level training reference rebuilt from the saved target-wise training OOD scores for the requested targets, weights, and row aggregate. For legacy fitted objects that do not contain those saved training scores, the original saved row-level reference is used when possible; otherwise NA.

  • targets.used: the number of weighted targets that contributed to each row-level score.

  • target.score: optional matrix of target-wise calibrated OOD scores, returned when return.details = TRUE.

  • target.delta: optional matrix of raw target-wise reconstruction discrepancies, returned when return.details = TRUE.

  • target.reconstruction: when return.reconstruction = TRUE, a data frame containing the saved learners' predictions for the scored targets.

  • reconstructed.data: when return.reconstruction = TRUE, the harmonized raw table with scored targets replaced by their reconstructions. Other columns retain their harmonized values. Integer restoration applies to generated reconstructions, preserving observed numeric values in the other columns.

  • completed.data: when both return.details and return.reconstruction are TRUE, the raw predictor table after prediction-time imputation. This is distinct from the target reconstruction table.

  • info: a list of diagnostics including harmonization details, dropped supervised response columns when present, row-wise unseen-factor flags, learner-loading information, supervised-auxiliary diagnostics when present, the active row aggregate and its arguments, whether the saved row-level calibration was used, and any target-specific issues.

Arguments

formula

An optional symbolic model description passed to impute for the initial training-data imputation when supervised.formula is not supplied. It follows the same guidelines as in impute. This formula does not specify the full-sweep learner bank, the OOD targets, or the supervised auxiliary forest. When supervised.formula is supplied, the raw predictor block is taken from the right-hand side of supervised.formula, and formula is not used.

data

Training data, converted to a plain data frame before processing. Matrices, tibbles, and data.table objects can be supplied. Column names must be unique and nonempty, each column must be a vector, and numeric values must be finite or missing. Bare vectors and nested matrix or list columns are not supported. Variables that are not real-valued are coerced to factors before fitting when possible; otherwise fitting stops with an error. Rows and columns that are entirely missing are dropped before training begins. If supervised.formula is supplied, data should contain both the supervised response columns and the raw predictor columns. The learned imputer is then built on the raw predictor block defined by the right-hand side of supervised.formula.

ntree, nodesize, nsplit, nimpute, fast, blocks, max.iter, ytry, always.use, verbose

Arguments passed to impute for the initial training-data imputation. The argument full.sweep is controlled internally and should not be supplied. fast is also forwarded to the saved target learners and, unless overridden by supervised.args, to the supervised forest. verbose controls progress throughout training, prediction, and scoring. The same verbose, max.predict.iter, eps, and cache.learners controls are also used by predict.impute.learn and impute.ood. Iteration and forest-size controls must be scalar integers in their permitted ranges; max.iter must be at least one.

mf.q

Controls the imputation engine used by impute. With mf.q = 1 and always.use = NULL, targets are updated one at a time using missForest. Other positive settings use the multivariate missForest generalization; a fraction below one controls the proportion of missing-data variables grouped as responses, and a value above one specifies their requested group size. A non-NULL always.use selects the multivariate branch, including an empty vector or a vector with no matching column names. If mf.q is omitted, training uses on-the-fly imputation. The selected method is recorded in manifest$train.imputation.

eps

Finite nonnegative convergence threshold. In impute.learn this controls the initial training-data imputation. In predict.impute.learn and impute.ood it controls early stopping for the prediction-time sweep.

...

For impute.learn, additional arguments passed to impute. For predict.impute.learn and impute.ood, additional arguments are currently ignored.

supervised.formula

Optional supervised learning formula used to augment the learned imputer for improved OOD detection in supervised settings. The left-hand side defines the supervised response and the right-hand side defines the raw predictor block to be learned by impute.learn. The supervised forest is fit internally after the raw predictor block has been completed. Its training out-of-bag predicted values and test-time predicted values are appended internally as auxiliary predictors. Training auxiliary values use out-of-bag predictions when available, with ordinary forest predictions and training means used as fallbacks. At deployment, auxiliary values are computed once from the initialized raw predictors and held fixed during the imputation passes. Predictors must be raw column names; non-syntactic names can be enclosed in backticks. Interactions and transformations are not supported. These auxiliary predictors are internal and are not expected in newdata.

supervised.args

Optional named list of arguments passed to the internal supervised rfsrc fit. Entries for formula, data, and forest are controlled internally and are ignored if supplied. All other entries override the defaults used for this supervised fit, including defaults inherited from top-level arguments such as fast and internally supplied defaults such as perf.type = "none".

full.sweep.options

A named list of options used when fitting saved target learners after initial training-data imputation. Defaults are ntree = 100, nodesize = NULL, and nsplit = 10, independently of the corresponding settings used for initial imputation. Recognized entries include ntree, nodesize, nsplit, mtry, splitrule, bootstrap, sampsize, samptype, perf.type, rfq, save.memory, importance, proximity, terminal.qualts, and terminal.quants. To request on-the-fly prediction for the saved target learners, supply terminal.qualts = FALSE and terminal.quants = FALSE in this list. These settings are passed to the fitting function selected by anonymous; they do not change the options for initial imputation or for the supervised auxiliary forest, which uses supervised.args. Unknown entries are ignored with a warning. Duplicate names produce a warning, and the last entry for each name is used.

target.mode

Determines which raw variables receive a saved full-sweep learner. The default "missing.only" saves learners only for variables that were missing in the training data. The option "all" saves a learner for every retained raw variable. If the training data are complete, target.mode = "all" must be used. When supervised.formula is supplied, the default is promoted internally to "all" so that every retained raw predictor can later be updated and scored. For the broadest OOD coverage, target.mode = "all" is recommended so every deployment-time variable can be reconstructed.

deployment.xvars

Controls which raw predictors are assumed to be available later when the saved imputer is used on new data. If NULL, all raw non-target columns are used. If a character vector, the same raw predictor set is used for all targets. If a named list, names must be target variables and each entry gives the raw predictor set for that target; targets omitted from the list fall back to all raw non-target columns. Unnamed or duplicated list entries are not allowed. Entries named for non-target variables, and predictor names not found in the training data, are ignored with a warning. When supervised.formula is supplied, the internal auxiliary predicted.* columns are added automatically on top of these raw predictor sets. In other words, deployment.xvars restricts the raw predictors only.

anonymous

If TRUE, uses rfsrc.anonymous when fitting the saved target-wise full-sweep learner bank. The internal supervised forest, when requested through supervised.formula, is fit with rfsrc because it must be retained for later prediction-time auxiliary variables. Thus anonymous usually reduces the size of the target-wise learner bank, but not necessarily the size of the supervised forest.

learner.prefix, learner.root

Names used when writing saved full-sweep learners to disk. If supervised.formula is supplied, the internal supervised forest is saved under the same learner root. learner.root must be a relative directory path within the imputer directory. learner.prefix must be a single file-name component. Absolute paths and parent-directory components are not allowed.

out.dir

Optional output directory. If supplied and save.on.fit = TRUE, the manifest and the saved full-sweep learners are written to this directory during fitting. This requires the fst package because learners are serialized with fast.save. If supervised.formula is supplied, the internal supervised forest is also saved there.

wipe

If TRUE, replaces an existing output directory after the new imputer has been written and verified. If FALSE, unrelated existing files are retained. Saving uses a temporary bundle and requires additional disk space.

keep.models

If TRUE, keeps the fitted full-sweep learners in memory in the returned object. At least one storage mode must be enabled: either keep.models = TRUE or out.dir with save.on.fit = TRUE. If supervised.formula is supplied, the internal supervised forest is kept in memory under the same rule.

keep.ximp

If TRUE, keeps the completed raw training predictor table in the returned object. This is not required for later prediction.

save.on.fit

If TRUE and out.dir is supplied, writes the imputer to disk during fitting.

save.ood

If TRUE, computes and stores an OOD reference. The reference is built during training from target-wise out-of-bag reconstruction discrepancies, their target-wise calibrated training scores, and a default row-level weighted mean using the saved OOD target weights. If no weight is supplied at fit time, equal target weights are used. The saved target-wise training-score matrix allows impute.ood to rebuild a calibrated row-level percentile for arbitrary target subsets, test-time weight overrides, and alternate row aggregates besides the weighted mean.

object

An object returned by impute.learn or load.impute.learn.

path

Directory containing a saved imputer. Use a dedicated directory rather than a filesystem root, home directory, working directory, or an ancestor of these. A complete bank can be saved back to its source path. Otherwise, source and destination directories must not contain one another. An object loaded with a target subset must be saved to a different directory. Save and load operations require the fst package because learners are read and written with fast.save and fast.load.

targets

Optional subset of target variables to load, update, or score. Unknown names are ignored with a warning. For impute.ood, row-level percentile calibration is rebuilt for the requested target subset from the saved target-wise training OOD scores whenever those scores are available in the manifest. A load-time subset limits which learners are available for predictor completion. It can therefore differ from scoring a subset of a fully loaded bank when other predictors are missing.

lazy

If TRUE, saved learners are loaded only when they are needed. If FALSE, all saved learners are loaded at once. If supervised.formula was used at fit time, the internal supervised forest follows the same lazy versus eager loading rule.

newdata

New data to be imputed or scored, converted to a plain data frame before processing. The column-name and vector-column requirements for data also apply. Missing columns are added and extra columns are dropped to match the training schema. A zero-row table returns correctly typed empty output without calling forests. Retained numeric values must be finite or missing. Nonmissing values that cannot be converted to a required numeric type become missing with a warning; their row indices are recorded in conversion.issues. Unseen factor levels are converted to NA for harmonization, but they are also tracked row-wise. In supervised mode, newdata should contain only the raw predictor columns learned by the imputer. The internal auxiliary columns are created automatically and should not be supplied by the user. If supervised response columns are supplied in newdata and are not part of the learned raw predictor block, they are treated as extra columns and dropped; they are not used to compute auxiliary predictors, imputations, or OOD scores. In impute.ood, observed raw target values are used to compute target-wise discrepancies. Missing raw target values are completed for predictor-side reconstruction but do not themselves contribute a finite target-wise OOD score. Any row containing an unseen factor level is flagged and its row-level OOD score is set to the maximum value.

max.predict.iter

Maximum number of full-sweep passes applied to newdata before the saved learner bank is used for OOD reconstruction or returned prediction-time imputations. Must be a nonnegative integer. Zero performs initialization without iterative target updates; OOD reconstruction is still performed.

restore.integer

If TRUE, generated imputations in integer-supported variables are rounded to the integer grid. This includes columns stored as R integers and numeric columns whose observed finite training values were all integer-valued. Observed numeric values in newdata are preserved, including fractional values, both as predictors during imputation and in the returned data. Restoration applies to initialization values and model-based updates. Numeric training columns retain numeric storage. Columns stored as R integers during training are returned as integer vectors only when every nonmissing completed value is exactly integral and within R's integer range; otherwise numeric storage preserves observed fractions and out-of-range values. Set to FALSE to retain unrounded generated numeric values. Factor columns are always conformed back to the training schema. The package operates on real-valued and factor variables; inputs that are not real-valued are coerced to factors during preprocessing when possible, otherwise an error is raised.

cache.learners

How saved learners are reused during prediction or OOD scoring. For predict.impute.learn, the default "session" loads each needed learner once per call. The option "none" reloads a learner every time it is needed. The option "all" loads all requested learners before work starts. For impute.ood, "all" is the default because the saved learner bank is typically reused once for predictor-side completion and again for target reconstruction.

weight

Optional nonnegative target weights used for row-level OOD aggregation. In impute.learn, these weights define the default row-level OOD weighting scheme stored in the manifest. In impute.ood, they define the active row-level weighting scheme for the current scoring call. If supplied as a named vector, entries are matched to targets by name; omitted targets are set to zero, and extra names are ignored. If omitted in impute.ood, the saved training-time OOD weights are used automatically. Because the fit stores target-wise training OOD scores, score.percentile can be recalibrated automatically for test-time weight overrides rather than being limited to the original training-time weights.

aggregate

Row-level aggregation metric used by impute.ood to combine calibrated target-wise OOD scores. The default "bounded.product" applies a weighted product of the form \(1-\prod_j \max(1-u_j,\varepsilon)^{\tilde w_j}\), where \(u_j\) is the calibrated score for target \(j\) and \(\tilde w_j\) denotes weights normalized over the positively weighted, scoreable targets in the row. "weighted.mean" is the weighted average. "weighted.lp" applies a weighted Minkowski \(L_p\) aggregation to the calibrated target scores. "weighted.lp.log" first applies the tail-stretching transform \(-\log\{\max(1-u_j,\varepsilon)\}\) to each calibrated target score \(u_j\) and then applies the weighted \(L_p\) aggregation. "top.k" averages only the \(k\) largest scoreable target scores among the positively weighted targets. When the fit stores target-wise training OOD scores, score.percentile is rebuilt for the requested aggregate as well as the requested targets and weights.

aggregate.args

Optional named list of tuning arguments for aggregate. Recognized entries are p for "weighted.lp" and "weighted.lp.log", k (or top.k) for "top.k", and eps for "weighted.lp.log" and "bounded.product". The default values are p = 2, k = 1, and eps = 1e-12. The power must be finite and at least one; the probability floor must satisfy \(0 < \varepsilon < 1\). The top-\(k\) value must be finite, at least one, and no larger than .Machine$integer.max; it is rounded to an integer. Unknown entries are ignored with a warning.

return.details

If TRUE, impute.ood returns the per-target discrepancy and calibrated target-score matrices and a richer info list.

return.reconstruction

If TRUE, impute.ood returns the target-wise reconstructed values used during OOD scoring. The reconstructed values are returned on the raw target scale, with one column per scored target. Typically these values are computed internally and discarded.

Author

Hemant Ishwaran and Udaya B. Kogalur

Details

Training

Training begins by converting variables that are not real-valued to factors when possible; otherwise fitting stops with an error. Rows and columns that are entirely missing are removed before the training schema is stored. The imputer is then fitted in two stages:

  1. Complete the training data. Missing values are imputed using impute, with the same options as that function. With mf.q = 1 and always.use = NULL, targets are updated one at a time. Other positive settings use the multivariate missForest generalization.

    If mf.q is omitted, on-the-fly imputation is used when formula is specified; otherwise default unsupervised imputation is used. Complete training data with target.mode = "all" skip this stage.

  2. Fit the saved learners. A full forest sweep is fitted on the completed training data. For each target selected by target.mode, a forest is fitted using rows where that target was originally observed and predictors selected by deployment.xvars. All saved learners use the same completed training table; fitting them does not further update that table.

Training stops if every requested target learner fails. If some succeed, a partial bank is returned with one summary warning. manifest$learners records each target's status and error, and printing the imputer reports successful and unavailable learners.

Choosing predictors

By default, deployment.xvars = NULL uses every non-target column as a predictor. Restrict deployment.xvars when the training data include outcomes, future-only variables, identifiers, or other fields that could introduce leakage or will be unavailable in new data.

Supervised auxiliary predictors

Supplying supervised.formula fits an internal supervised forest. Its out-of-bag predicted values for training data and predicted values for new data provide auxiliary predictors for the saved learners. These variables can influence both imputation and OOD scoring, but are not imputation targets. Leave supervised.formula unspecified to use the basic unsupervised learned imputer.

The auxiliary variables are created automatically and added to each saved target's predictor set. deployment.xvars restricts only the raw predictors; users do not need to supply auxiliary columns in new data.

Saving and loading

A saved imputer consists of a small manifest and a directory of learners. Each learner is saved with fast.save and loaded with fast.load, so save and load operations require the fst package. The supervised forest, when used, is saved alongside the target learners. The explicit save method can save learners from memory or load them from an attached saved path.

Both training-time and explicit saving write and verify the new learners before replacing an existing destination. A failed staging operation leaves the previous saved imputer unchanged. If replacement fails, restoration of the previous directory is attempted; an unrecovered backup path is reported.

Imputing new data

Prediction applies the saved learners in three steps:

  1. Match the training schema. Columns and types in newdata are matched to the training schema. Supervised response columns outside the learned raw predictor block are dropped and are not used for prediction or OOD scoring.

  2. Initialize missing values. Missing raw values are filled with training means or modes. In supervised mode, the saved supervised forest then computes auxiliary predictors from this initialized table. The auxiliary predictors stay fixed throughout the subsequent passes.

  3. Update with saved learners. Full-sweep passes update missing values in the selected targets. Each target update requires one valid prediction per requested row. Failed or unavailable predictions leave the previous imputed values unchanged and are recorded in target.issues. A pass with no valid model updates is reported separately from convergence.

With target.mode = "missing.only", a variable that was complete in training but is missing in new data receives only an initialization value. Use target.mode = "all" when missing values may appear later in any raw variable. Complete training data also require this setting because there are no missing variables from which to select the saved targets.

Integer restoration applies only to generated values, including initialization values in columns without a target update. Observed numeric entries are unchanged.

OOD training references

With save.ood = TRUE, each saved learner's out-of-bag predictions are compared with observed training values to form target-wise reconstruction discrepancies:

  • Continuous and integer targets use absolute reconstruction error.

  • Factor targets use negative log predictive probabilities. Unavailable or invalid probabilities give missing discrepancies; zero probabilities are scored using the probability floor.

The references are stored in the manifest. Each learner entry records n.oob.finite and n.oob.nonfinite.

The training-time row reference combines calibrated target scores using a weighted mean. If weight is omitted, all saved OOD targets receive weight 1. Named weights are matched to targets, and omitted targets receive weight 0. These weights are also the defaults for later OOD scoring.

OOD scoring

impute.ood first completes the predictors in newdata using the same schema matching, initialization, and full-sweep passes as predict.impute.learn. It then reconstructs each requested raw target from its saved learner and compares the reconstruction with the observed value in newdata. A target that is missing in a row does not contribute to that row's OOD score.

Discrepancies are converted to target-wise OOD scores using the saved target-specific training references. Values strictly above a nonempty reference's maximum, including positive infinity, receive its largest stored probability. Ties, including equality at the largest quantile, follow the reference's quantile-grid convention. Missing discrepancies and empty references are unscored.

In supervised mode, the auxiliary variables are predictors for reconstruction; the supervised response is not used in new data.

Row-level OOD scores

impute.ood returns two row-level summaries:

score

Combines calibrated target scores over the targets that are observed and scoreable in each row. The default is a bounded product rule. Alternatives are a weighted mean, weighted \(L_p\), log-tail weighted \(L_p\), and top-\(k\) rules. These options allow greater sensitivity to sparse but severe coordinate shifts. Unless overridden, scoring uses the weights saved by impute.learn.

score.percentile

Calibrates the row score against a training reference rebuilt from the saved target-wise training OOD scores. The reference uses the requested target subset, weights, and row aggregate, so percentile calibration remains available when any of these settings change.

Unseen factor levels

Unseen factor levels are tracked by row when matching new data to the training schema. impute.ood flags these rows and assigns them the maximum row-level score. If an unseen level occurs in a scored target, its target-level discrepancy is also maximal.

References

Stekhoven D.J. and Buhlmann P. (2012). MissForest--non-parametric missing value imputation for mixed-type data. Bioinformatics, 28(1):112--118.

Tang F. and Ishwaran H. (2017). Random forest missing data algorithms. Statistical Analysis and Data Mining, 10:363--377.

See Also

impute.rfsrc, rfsrc, predict.rfsrc.

Examples

Run this code
## ------------------------------------------------------------
## small data example: uses missForest for impute engine
## ------------------------------------------------------------

set.seed(101)
aq <- airquality[, c("Ozone", "Solar.R", "Wind", "Temp", "Month")]
aq$Month <- factor(aq$Month)

id <- sample(1:nrow(aq), 100)
train <- aq[id, ]
test <- aq[-id, ]

## training the imputer
fit <- impute.learn(
  data = train,
  ntree = 25,
  mf.q = 1,
  max.iter = 5,
  full.sweep.options = list(ntree = 25, nsplit = 5)
)

## test time imputation
test.imp <- predict(fit, test, max.predict.iter = 2, verbose = FALSE)
print(head(test.imp))

# \donttest{

## OOD scoring is most informative when every deployment-time
## variable can be reconstructed, so target.mode = "all" is recommended.
## Optional named OOD weights can also be supplied here. Any omitted
## targets receive weight 0, and the saved weights are reused
## automatically later by impute.ood().

ood.fit <- impute.learn(
  data = train,
  ntree = 25,
  mf.q = 1,
  max.iter = 5,
  target.mode = "all",
  save.ood = TRUE,
  full.sweep.options = list(ntree = 25, nsplit = 5),
  verbose = FALSE
)

ood <- impute.ood(ood.fit, test, return.details = TRUE, verbose = FALSE)
print(head(ood$score))
print(head(ood$score.percentile))

## try a more spike-sensitive row aggregate
ood.lp <- impute.ood(ood.fit, test,
                     aggregate = "weighted.lp",
                     aggregate.args = list(p = 4),
                     verbose = FALSE)
print(head(ood.lp$score.percentile))

## ------------------------------------------------------------
## supervised OOD example: regression benchmark
## the user supplies only the raw x-columns at test time
## ------------------------------------------------------------

friedman1_sim <- function(n = 150, p = 10, sigma = 1) {
  X <- matrix(runif(n * p), nrow = n)
  y <- 10 * sin(pi * X[, 1] * X[, 2]) +
       20 * (X[, 3] - 0.5)^2 +
       10 * X[, 4] + 5 * X[, 5] +
       rnorm(n, sd = sigma)
  list(X = X, y = y)
}


trn <- data.frame(friedman1_sim())
tst <- data.frame(friedman1_sim())
xvars <- setdiff(names(trn), "y")


## impute data using missForest, construct a supervised forest
## - supervised forests are used to create auxiliary variables
## - improves test time OOD in supervised problems

sup.fit <- impute.learn(
  data = trn,
  mf.q = 1,
  supervised.formula = y ~ .,
  supervised.args = list(ntree = 50, nsplit = 5),
  full.sweep.options = list(ntree = 25, nsplit = 5),
  save.ood = TRUE,
  verbose = FALSE
)

## add some missing values to the test data
xnew <- tst[, xvars, drop = FALSE]
xnew[sample(seq_len(nrow(xnew)), 5), xvars[1]] <- NA
xnew[sample(seq_len(nrow(xnew)), 5), xvars[2]] <- NA

## imputation
xnew.imp <- predict(sup.fit, xnew, max.predict.iter = 2, verbose = FALSE)
print(head(xnew.imp))

## OOD score
ood.sup <- impute.ood(sup.fit, xnew, verbose = FALSE)
print(head(ood.sup$score.percentile))

## ------------------------------------------------------------
## Save the learned imputer to disk and load it later.
## This explicit save example writes learners kept in memory.
## Uses missForest for the impute engine.
## ------------------------------------------------------------

bundle.dir <- file.path(tempdir(), "aq.imputer")

fit <- impute.learn(
  data = train,
  ntree = 25,
  mf.q = 1,
  max.iter = 5,
  full.sweep.options = list(ntree = 25, nsplit = 5),
  keep.models = TRUE,
  verbose = FALSE
)

save.impute.learn(fit, bundle.dir, verbose = FALSE)
imp <- load.impute.learn(bundle.dir, lazy = TRUE, verbose = FALSE)
test.imp <- predict(imp, test, max.predict.iter = 2, verbose = FALSE)

unlink(bundle.dir, recursive = TRUE)



## ------------------------------------------------------------
## Challenging example with factors, uses save/reload
## ------------------------------------------------------------

## load pbc, convert everything to factors
data(pbc, package = "randomForestSRC")
dta <- data.frame(lapply(pbc, factor))
dta$days <- pbc$days
dta$status <- dta$status

## split the data into unbalanced train/test data (25/75)
## the train/test data have the same levels, but different labels
idx <- sample(1:nrow(dta), round(nrow(dta) * .25))
train <- dta[idx,]
test <- dta[-idx,]

## even harder ... factor level not previously encountered in training
levels(test$stage) <- c(levels(test$stage), "fake")
test$stage[sample(seq_len(nrow(test)), 10)] <- "fake"

## train forest
fit <- suppressWarnings(
  impute.learn(Surv(days, status) ~ ., train,
               target.mode = "all",
               save.ood = TRUE,
               keep.models = TRUE)
)

## save/reload
bundle.dir <- file.path(tempdir(), "pbc.imputer")
save.impute.learn(fit, bundle.dir, verbose = FALSE)
imp <- load.impute.learn(bundle.dir, lazy = TRUE, verbose = FALSE)
test.imp <- predict(imp, test, max.predict.iter = 2, verbose = FALSE)
ood <- impute.ood(imp, test, return.details = TRUE, verbose = FALSE)
print(which(ood$info$unseen.rows))
print(summary(test.imp))
unlink(bundle.dir, recursive = TRUE)
# }

Run the code above in your browser using DataLab