This function reads in a Tucson (decadal) format file of ring widths (.rwl).
read.tucson(fname, header = NULL, long = FALSE,
encoding = getOption("encoding"),
edge.zeros = TRUE, verbose = TRUE,
comment.char = "#", fix.dup.char = "X",
fill.internal.NA = NULL, strict = FALSE)An object of class c("rwl", "data.frame") with the series in
columns and the years as rows. The series IDs are the
column names and the years are the row names. Columns are returned in
the order the series appear in the file.
The returned object carries a "dplR.provenance" attribute recording
what the reader saw: the header lines, the precision each series was measured
at and whether the file mixes them, any series the reader renamed (as
old and new), the interior gaps and what the file held at
each, and one row per recoverable problem found while parsing. Most of this
cannot be recovered from the returned data or by reading the file again, and
rwl.check reports on it. Subsetting an rwl object with
[.rwl carries the record along and cuts it down to the series
and years that are left, but functions that return a new object, such as
detrend, drop it, so it describes the read rather than
travelling with the data indefinitely.
a character vector giving the file name of the
rwl file.
ignored, and accepted only so that existing calls do not fail. Header lines are now identified per line by their content. See ‘Details’.
ignored, and accepted only so that existing calls do not
fail. The two possible column layouts are now detected per line, so
8-character series IDs and years before -999 may occur in the
same file. As in read.crn, long here means a wider
fixed-width year field; it is unrelated to the “long” or
“tidy” layout of read.sheet, which is selected with
layout = "long". See ‘Details’.
the encoding of the file, used only if the file is not
valid UTF-8. Any name iconv accepts will do,
e.g. "latin1" or "ISO-8859-2". The default,
getOption("encoding"), means no encoding has been declared: the
reader then falls back to "latin1" and says so. See
‘Encoding’ below.
logical flag indicating whether leading or
trailing zeros in series will be preserved (when the flag is
TRUE, the default) or discarded, i.e. marked as NA
(when FALSE).
logical flag, print info on data. When
TRUE (the default) every interior gap is listed with the series,
the years, and what the file actually held there.
a character giving the character that marks a
comment line. Lines containing it are dropped.
a character appended to a series ID
when the file contains more than one record under that ID and
the records overlap in time. Appended repeatedly if necessary, so that
no measurement is lost and no name collides.
what to do with interior gaps, i.e. years inside
a series for which the file records no measurement. The default,
NULL, leaves them as NA. Any other value is passed to
fill.internal.NA, so 0 reproduces the output of
read.tucson.legacy, and "Mean", "Spline" or
"Linear" interpolate. See ‘Details’.
logical flag. When FALSE, the default, every
recoverable problem is reported as a warning and reading continues.
When TRUE the same problems become errors, so that a pipeline can
refuse a file rather than carry a warned guess forward.
Ring widths are digits, but the header lines and occasionally the series
IDs are not, and archived files carry accented site names and
investigator names in whatever 8-bit encoding the contributor's machine
used. read.tucson resolves this in three steps.
First, the file is tested against UTF-8. This is a decision and not a guess: UTF-8 is self-validating, and plain ASCII passes, so virtually every file takes this path and nothing is reported.
If the file is not valid UTF-8 and encoding was supplied,
that encoding is used and an ENCODING_DECLARED event is recorded.
If the declared encoding does not fit the file, read.tucson stops
rather than transcode into replacement characters.
If the file is not valid UTF-8 and no encoding was supplied, it
is read as "latin1" and an ENCODING_ASSUMED event is
recorded, with a warning naming the line and quoting the text the
assumption produced, so the guess can be judged at a glance. Under
strict = TRUE this becomes an error.
The fallback guesses rather than running a charset detector because
detectors do not work on this kind of file. They are byte-frequency
language models and need a volume of non-ASCII text that a file
of ring widths never has; on a rwl that is ASCII apart from one
accented name, ICU's detector scores the right answer at 0.16
and ranks UTF-16 alongside it. "latin1" is used instead
because it is total (every byte decodes, so it cannot fail) and lossless
(the round trip is byte-identical, so a wrong guess discards nothing and
the file can simply be re-read with encoding set).
Original reader by Hung Nguyen. Reworked and extended by Andy Bunn.
This reads in a standard rwl file as defined according to the standards of the ITRDB at https://www.ncei.noaa.gov/pub/data/paleo/treering/treeinfo.txt.
This function was replaced in dplR 1.8.0. The previous implementation is
still available, unchanged, as read.tucson.legacy. The two
differ in at least three important ways.
Interior gaps are no longer filled with zero. This is the change
most likely to alter your numbers. Where a file records no measurement
for a stretch of years inside a series, the old reader silently wrote
zeros. A ring width of zero means a locally absent ring, which is a real
observation about a tree in a year; a gap means the ring was not measured,
which is a statement about the data. These are different claims and the
reader no longer substitutes one for the other. Set
fill.internal.NA = 0 to recover the old values exactly.
Note that several dplR functions, detrend among them, do not
accept NA inside a series and will fail on such data. If you need
to detrend a series containing interior gaps you must decide what those
gaps mean and fill them deliberately, with fill.internal.NA
or with fill.internal.NA = 0 here.
Malformed files are reported rather than guessed at. Where the ten
fixed-width columns and a whitespace split of the same line disagree, the
line does not conform to the format and nothing in it says which reading
was meant, so it is returned as NA with both readings reported.
Repeated series IDs are reported instead of being silently
welded into one series, and are renamed when their records overlap. Tab
characters, which have no defined width in a fixed-width format, are
expanded to 8-column tab stops and reported.
The header and column layout are determined per line. This is why
header and long are no longer used. header could
only skip a fixed three lines and so could not describe the 4- and 6-line
headers that real files contain. long is worse: years before -999
need five columns and so take column 8 away from the series ID,
but that varies line by line, so a file mixing 8-character IDs
with BC dates is read wrongly whichever way a per-file switch is
set. Both are now decided from the content of each line. The arguments
are still accepted, in their original positions, so that existing code and
read.rwl's pass-through of ... do not fail with an
“unused argument” error; passing either one warns.
Editorial: A closing word on the format itself. The Tucson decadal format
dates from the era of the punch card. It is nominally fixed-width,
yet files in the ITRDB arrive with tab characters sitting in the
measurement fields, and a tab has no defined width at all. Column 8
belongs to the series ID unless the year is earlier than -999, in
which case the year claims it, and nothing but a minus sign distinguishes
the two readings. A negative number marks missing data, except when it
marks the end of a record; 999 ends a record in a 0.01 mm file, so a
genuine 9.99 mm ring cannot be told apart from a terminator. A single
ID may carry several records written in any order at all: in
ITRDB ut542.rwl a series runs to 1841 and then starts over at
1353 further
down the file. A decade line may resume off its boundary and quietly
swallow the years it skipped, as chil012.rwl does at 1242 for series
AGU034. None of
this is the fault of the people who made these data available, and the
tree-ring community deserves credit for sharing its measurements
openly long before most of science thought to do so. But we are still
paying for a format built around eighty columns of cardboard.
read.tucson tries to read
what the file says, to report plainly what it did, and to refuse
rather than guess when a line admits two readings. Even so, and despite
our best efforts, this format can and does fail.
read.tucson.legacy, read.rwl,
read.compact, read.tridas,
read.fh, write.tucson,
fill.internal.NA, [.rwl
## A short file with an interior gap: the 1910s are marked -999,
## i.e. not measured.
tf <- tempfile()
writeLines(c("TST01A 1900 123 134 145 156 167 178 189 150 141 132",
"TST01A 1910 -999 -999 -999 -999 -999 -999 -999 -999 -999 -999",
"TST01A 1920 211 222 233 244 255 266 277 288 299 300",
"TST01A 1930 311 999"), tf)
## The default reports the gap and returns it as NA.
x <- read.tucson(tf)
x["1915", ]
## Ask for the old behaviour explicitly.
x0 <- read.tucson(tf, fill.internal.NA = 0)
x0["1915", ]
unlink(tf)
Run the code above in your browser using DataLab