Learn R Programming

DCC (version 1.2.1)

dcc_detect_chunked: Run record-local checks over a file in chunks

Description

Larger-than-memory detection with an adaptive backend: chunks (or Arrow record batches) of chunk_size rows are read and checked one at a time, so peak memory is bounded by the chunk. Findings are identical to an in-memory dcc_detect for record-local checks (range, set, missing_items, straightlining, trap_items, and response_time with the median-relative cut disabled via min_median_ratio: ~). Cross-record checks (score_anomaly, median-relative response time) are rejected with typed errors; expr rules must not use aggregate functions.

Usage

dcc_detect_chunked(path, rules, chunk_size = 100000L, id_var = NULL,
  sep = NULL, backend = c("auto", "csv", "arrow"), encoding = "auto")

Value

A dcc_findings table with n_rows (total rows scanned), n_chunks, and backend attributes.

Arguments

path

Path to the input file: a delimited text file (CSV/TSV) for the csv backend, or a Parquet/Feather file for the arrow backend.

rules

A dcc_ruleset from dcc_rules.

chunk_size

Rows per chunk / record batch (default 100000).

id_var

Name of the record-id column, or NULL to use global row numbers (consistent across chunks).

sep

Field separator for the csv backend. NULL (default) infers the separator from the extension -- a tab for .tsv, a comma otherwise; pass an explicit value to override. Ignored by the arrow backend.

backend

One of "auto" (default), "csv", or "arrow".

encoding

Encoding of the csv backend input: "auto" (default, auto-detected) or an explicit name such as "UTF-8" or "latin1". Pass one explicitly when auto-detection misfires on pure-ASCII or low-signal data. Ignored by the arrow backend.

Details

Two backends share this entry point and produce identical findings: the csv backend streams a delimited file with fread (fread-native UTF-8/latin1 encoding only; column types are locked from the first chunk so later chunks cannot drift; each record must lie on a single line, so embedded newlines in quoted fields are unsupported -- read such files whole with dcc_read), and the arrow backend streams a Parquet/Feather file as record batches (requires the arrow package; types come from the file schema and the columnar format is always UTF-8, so the encoding restriction does not apply). With backend = "auto" the backend is chosen from the file extension: arrow for .parquet/.feather, csv for .csv/.tsv/.txt.

See Also

dcc_detect for in-memory detection.

Examples

Run this code
csv <- tempfile(fileext = ".csv")
writeLines(c("sid,score", "S1,90", "S2,150", "S3,70"), csv)
rules_file <- tempfile(fileext = ".yaml")
writeLines(c(
  "checks:",
  "  - id: R001",
  "    type: range",
  "    variable: score",
  "    min: 0",
  "    max: 100"
), rules_file)
if (requireNamespace("yaml", quietly = TRUE)) {
  # stream the file two rows at a time; encoding set explicitly since
  # short ASCII files defeat charset auto-detection
  dcc_detect_chunked(csv, dcc_rules(rules_file), chunk_size = 2L,
                     id_var = "sid", encoding = "UTF-8")
}

Run the code above in your browser using DataLab