Selects a subset of n_max rows from a design X, meant to be
used as a cheap pre-fit reduction for large designs: fit on
X[idx, ]/y[idx] instead of the full data. Unlike
Vecchia/Nystrom (which still use every point), this discards
n - n_max points outright, in exchange for an ordinary exact fit
on the reduced design.
Sorted 1-based row-indices into X (and the matching
y) to keep.
Arguments
X
n x d design matrix.
n_max
Target subset size; if n_max >= nrow(X), returns all
indices (no-op).
method
"kmeans" (default): n_max k-means centroids on
X, each replaced by its nearest actual data point, so the
subset always consists of real observations; falls back to
"random" if k-means degenerates. "random": uniform
subsample without replacement.
seed
RNG seed (k-means initialization and/or random fallback).