Run a battery of integrity checks against a rwl object or a Tucson
format file and return the result as data: one row per finding, each carrying
a stable check ID and a severity. Intended both for a researcher
looking at one collection and for sweeping a large archive.
rwl.check(x, file = NULL,
checks = c("structure", "series", "values", "zeros",
"crossdating", "provenance", "file"),
control = rwl.check.control(), ...)rwl.check.control(min.depth = 2, depth.run = 10, min.length = 30,
run.min = 6,
lag.max = 5, min.overlap = 50, spline.nyrs = 32,
r.dating = 0.35,
r.margin = 0.05, outlier.mad = 4, outlier.gap = 0.2,
r.cohesion = 0.35,
plausible.mean = c(0.05, 10),
small.thresh = NA, big.thresh = NA,
max.year = as.integer(format(Sys.Date(), "%Y")))
rwl.check.catalogue()
An object of class "rwl.check": a list with the file name, the
control values used, a findings
data.frame, and a meta
list of summary values.
as.data.frame returns the findings, one row per finding, with columns
file, check, severity, group, series,
year.from, year.to, n, value and message.
summary returns a one-row data.frame holding the collection's
summary values, counts of errors, warnings and notes, and one count column per
check ID. Rows from different files have identical columns and can
be combined with rbind into a triage table.
a rwl object, an object coercible to one via
as.rwl, or a length-one character giving the path to a
Tucson format file, which is read with read.tucson.
a character naming the file the object came from. Used to
label findings and to run the "file" checks. Taken from x when
x is a path.
a character vector of check groups to run. The
"crossdating" group is much the most expensive; omit it for a fast
pass over many files.
a list of tuning values, from
rwl.check.control.
years measured by fewer than this many series are flagged.
how many consecutive such years are needed before they are reported. Every collection thins at its edges, so a year or two says nothing.
series with fewer than this many rings are flagged.
a run of this many identical non-zero rings is flagged.
largest lag searched when looking for dating errors.
rings a series must share with the master to be checked against it.
stiffness, in years, of the spline used to high pass the
series before the crossdating checks. Crossdating works on year-to-year
variation, so correlating raw ring widths would measure the agreement of
age trends instead. The default is the COFECHA value, which is
also caps's own default and what corr.rwl.seg
recommends.
correlation a series must reach at its best lag before a non-zero lag is called a dating error.
how much better than the correlation as dated that best-lag correlation must be.
how many median absolute deviations below its own collection's median correlation a series must fall to be flagged as not fitting the collection.
how far below the collection median, in correlation, that series must also fall. The second condition stops a very tight collection from flagging trivial spread.
median series correlation below which the collection as a whole is reported as sharing little common signal.
range, in mm, within which a collection's mean ring width is taken to be plausible.
thresholds for flagging small and large
rings. NA omits the check.
years after this are flagged.
not used.
Andy Bunn
rwl.check complements rwl.report. Where
rwl.report describes a collection for a reader, rwl.check looks
for defects and returns them as rows, so that examining one file and sweeping
ten thousand are the same operation.
Every check runs independently and none of them stops. A check that fails
becomes a finding with the ID RWL_CHECK_ERROR rather than an
error, so a pathological file still yields a report. This matters at scale:
the files that break are the ones worth looking at.
Each finding carries a stable check ID and one of three severities:
"error" for something that cannot be right, "warning" for
something probably wrong, and "note" for something worth knowing.
rwl.check.catalogue returns the full table of IDs, groups,
severities and descriptions. The IDs are a contract: they are what a
sweep filters on and what two runs are compared by, and they do not change
meaning between releases.
The "provenance" checks read the record read.tucson
attaches to what it returns, and report things that cannot be recovered any
other way. Chief among them is RWL_ID_RENAMED: where a file gives two
different cores the same id, the reader renames one, so a series on the
returned object carries a name that appears nowhere in the file. Nothing but
the reader can say that. The group is silent for an object that carries no
such record, including one read by read.tucson.legacy.
The "file" checks need the file itself and are skipped when only an
object is given. They are the only way to see line endings, stray tabs, and
the span declared in the ITRDB header, all of which are gone once
the file has been parsed into a matrix of numbers.
The crossdating checks ask whether a series fits the collection it is in
rather than whether it clears a fixed correlation, because what counts as a
good correlation is a property of the site: a high-elevation conifer stand
runs 0.7 to 0.8, while some collections, particularly ecological ones sampled
for growth rather than for a climate signal, sit far below that and are
perfectly sound. A collection with little common signal is therefore reported
once, as RWL_WEAK_COLLECTION and as a note rather than a warning: it
is not a fault to be corrected, but it does decide whether the series can be
crossdated against each other or carry a chronology, so it is worth saying.
The defaults in rwl.check.control were calibrated against the
rwl objects shipped with dplR. In particular RWL_DATING_LAG
fires on none of anos1, ca533,
co021, gp.rwl, nm046 or
wa082, and does fire on a series displaced by two years.
The value of an RWL_DATING_LAG finding is the lag at which
the series correlates best with the master, with the sign used by
ccf.series.rwl (with series.x = FALSE) and by
best.lag in corr.rwl.seg: negative means the series
is probably missing rings, positive that it probably has false rings or a
ring counted twice.
rwl.report, read.tucson,
interseries.cor, corr.rwl.seg
library(utils)
data(zof.rwl)
res <- rwl.check(zof.rwl, file = "zof.rwl")
res
## the findings as data
head(as.data.frame(res))
## one row per file: rbind these over an archive to get a work queue
data(ca533)
sweep <- rbind(summary(rwl.check(zof.rwl, file = "zof.rwl")),
summary(rwl.check(ca533, file = "ca533")))
sweep[, c("file", "n.series", "n.error", "n.warning", "n.note")]
## what is looked for
head(rwl.check.catalogue())
## a fast pass, skipping the crossdating checks
rwl.check(ca533, checks = c("structure", "series", "values", "zeros"))
Run the code above in your browser using DataLab