Learn R Programming

rlibkriging (version 1.2-3)

subsetOfData: Subset-of-data pre-fit reduction.

Description

Selects a subset of n_max rows from a design X, meant to be used as a cheap pre-fit reduction for large designs: fit on X[idx, ]/y[idx] instead of the full data. Unlike Vecchia/Nystrom (which still use every point), this discards n - n_max points outright, in exchange for an ordinary exact fit on the reduced design.

Usage

subsetOfData(X, n_max, method = "kmeans", seed = 123)

Value

Sorted 1-based row-indices into X (and the matching

y) to keep.

Arguments

X

n x d design matrix.

n_max

Target subset size; if n_max >= nrow(X), returns all indices (no-op).

method

"kmeans" (default): n_max k-means centroids on X, each replaced by its nearest actual data point, so the subset always consists of real observations; falls back to "random" if k-means degenerates. "random": uniform subsample without replacement.

seed

RNG seed (k-means initialization and/or random fallback).

Author

Yann Richet [email protected]

Examples

Run this code
X <- matrix(runif(200), ncol = 2)
idx <- subsetOfData(X, 20)
Xr <- X[idx, ]

Run the code above in your browser using DataLab