Learn R Programming

dplR (version 1.8.0)

read.tucson: Read Tucson Format Ring Width File

Description

This function reads in a Tucson (decadal) format file of ring widths (.rwl).

Usage

read.tucson(fname, header = NULL, long = FALSE,
            encoding = getOption("encoding"),
            edge.zeros = TRUE, verbose = TRUE,
            comment.char = "#", fix.dup.char = "X",
            fill.internal.NA = NULL, strict = FALSE)

Value

An object of class c("rwl", "data.frame") with the series in columns and the years as rows. The series IDs are the column names and the years are the row names. Columns are returned in the order the series appear in the file.

The returned object carries a "dplR.provenance" attribute recording what the reader saw: the header lines, the precision each series was measured at and whether the file mixes them, any series the reader renamed (as

old and new), the interior gaps and what the file held at each, and one row per recoverable problem found while parsing. Most of this cannot be recovered from the returned data or by reading the file again, and

rwl.check reports on it. Subsetting an rwl object with

[.rwl carries the record along and cuts it down to the series and years that are left, but functions that return a new object, such as

detrend, drop it, so it describes the read rather than travelling with the data indefinitely.

Arguments

fname

a character vector giving the file name of the rwl file.

header

ignored, and accepted only so that existing calls do not fail. Header lines are now identified per line by their content. See ‘Details’.

long

ignored, and accepted only so that existing calls do not fail. The two possible column layouts are now detected per line, so 8-character series IDs and years before -999 may occur in the same file. As in read.crn, long here means a wider fixed-width year field; it is unrelated to the “long” or “tidy” layout of read.sheet, which is selected with layout = "long". See ‘Details’.

encoding

the encoding of the file, used only if the file is not valid UTF-8. Any name iconv accepts will do, e.g. "latin1" or "ISO-8859-2". The default, getOption("encoding"), means no encoding has been declared: the reader then falls back to "latin1" and says so. See ‘Encoding’ below.

edge.zeros

logical flag indicating whether leading or trailing zeros in series will be preserved (when the flag is TRUE, the default) or discarded, i.e. marked as NA (when FALSE).

verbose

logical flag, print info on data. When TRUE (the default) every interior gap is listed with the series, the years, and what the file actually held there.

comment.char

a character giving the character that marks a comment line. Lines containing it are dropped.

fix.dup.char

a character appended to a series ID when the file contains more than one record under that ID and the records overlap in time. Appended repeatedly if necessary, so that no measurement is lost and no name collides.

fill.internal.NA

what to do with interior gaps, i.e. years inside a series for which the file records no measurement. The default, NULL, leaves them as NA. Any other value is passed to fill.internal.NA, so 0 reproduces the output of read.tucson.legacy, and "Mean", "Spline" or "Linear" interpolate. See ‘Details’.

strict

logical flag. When FALSE, the default, every recoverable problem is reported as a warning and reading continues. When TRUE the same problems become errors, so that a pipeline can refuse a file rather than carry a warned guess forward.

Encoding

Ring widths are digits, but the header lines and occasionally the series IDs are not, and archived files carry accented site names and investigator names in whatever 8-bit encoding the contributor's machine used. read.tucson resolves this in three steps.

First, the file is tested against UTF-8. This is a decision and not a guess: UTF-8 is self-validating, and plain ASCII passes, so virtually every file takes this path and nothing is reported.

If the file is not valid UTF-8 and encoding was supplied, that encoding is used and an ENCODING_DECLARED event is recorded. If the declared encoding does not fit the file, read.tucson stops rather than transcode into replacement characters.

If the file is not valid UTF-8 and no encoding was supplied, it is read as "latin1" and an ENCODING_ASSUMED event is recorded, with a warning naming the line and quoting the text the assumption produced, so the guess can be judged at a glance. Under strict = TRUE this becomes an error.

The fallback guesses rather than running a charset detector because detectors do not work on this kind of file. They are byte-frequency language models and need a volume of non-ASCII text that a file of ring widths never has; on a rwl that is ASCII apart from one accented name, ICU's detector scores the right answer at 0.16 and ranks UTF-16 alongside it. "latin1" is used instead because it is total (every byte decodes, so it cannot fail) and lossless (the round trip is byte-identical, so a wrong guess discards nothing and the file can simply be re-read with encoding set).

Author

Original reader by Hung Nguyen. Reworked and extended by Andy Bunn.

Details

This reads in a standard rwl file as defined according to the standards of the ITRDB at https://www.ncei.noaa.gov/pub/data/paleo/treering/treeinfo.txt.

This function was replaced in dplR 1.8.0. The previous implementation is still available, unchanged, as read.tucson.legacy. The two differ in at least three important ways.

Interior gaps are no longer filled with zero. This is the change most likely to alter your numbers. Where a file records no measurement for a stretch of years inside a series, the old reader silently wrote zeros. A ring width of zero means a locally absent ring, which is a real observation about a tree in a year; a gap means the ring was not measured, which is a statement about the data. These are different claims and the reader no longer substitutes one for the other. Set fill.internal.NA = 0 to recover the old values exactly.

Note that several dplR functions, detrend among them, do not accept NA inside a series and will fail on such data. If you need to detrend a series containing interior gaps you must decide what those gaps mean and fill them deliberately, with fill.internal.NA or with fill.internal.NA = 0 here.

Malformed files are reported rather than guessed at. Where the ten fixed-width columns and a whitespace split of the same line disagree, the line does not conform to the format and nothing in it says which reading was meant, so it is returned as NA with both readings reported. Repeated series IDs are reported instead of being silently welded into one series, and are renamed when their records overlap. Tab characters, which have no defined width in a fixed-width format, are expanded to 8-column tab stops and reported.

The header and column layout are determined per line. This is why header and long are no longer used. header could only skip a fixed three lines and so could not describe the 4- and 6-line headers that real files contain. long is worse: years before -999 need five columns and so take column 8 away from the series ID, but that varies line by line, so a file mixing 8-character IDs with BC dates is read wrongly whichever way a per-file switch is set. Both are now decided from the content of each line. The arguments are still accepted, in their original positions, so that existing code and read.rwl's pass-through of ... do not fail with an “unused argument” error; passing either one warns.

Editorial: A closing word on the format itself. The Tucson decadal format dates from the era of the punch card. It is nominally fixed-width, yet files in the ITRDB arrive with tab characters sitting in the measurement fields, and a tab has no defined width at all. Column 8 belongs to the series ID unless the year is earlier than -999, in which case the year claims it, and nothing but a minus sign distinguishes the two readings. A negative number marks missing data, except when it marks the end of a record; 999 ends a record in a 0.01 mm file, so a genuine 9.99 mm ring cannot be told apart from a terminator. A single ID may carry several records written in any order at all: in ITRDB ut542.rwl a series runs to 1841 and then starts over at 1353 further down the file. A decade line may resume off its boundary and quietly swallow the years it skipped, as chil012.rwl does at 1242 for series AGU034. None of this is the fault of the people who made these data available, and the tree-ring community deserves credit for sharing its measurements openly long before most of science thought to do so. But we are still paying for a format built around eighty columns of cardboard. read.tucson tries to read what the file says, to report plainly what it did, and to refuse rather than guess when a line admits two readings. Even so, and despite our best efforts, this format can and does fail.

See Also

read.tucson.legacy, read.rwl, read.compact, read.tridas, read.fh, write.tucson, fill.internal.NA, [.rwl

Examples

Run this code
## A short file with an interior gap: the 1910s are marked -999,
## i.e. not measured.
tf <- tempfile()
writeLines(c("TST01A  1900   123   134   145   156   167   178   189   150   141   132",
             "TST01A  1910  -999  -999  -999  -999  -999  -999  -999  -999  -999  -999",
             "TST01A  1920   211   222   233   244   255   266   277   288   299   300",
             "TST01A  1930   311   999"), tf)

## The default reports the gap and returns it as NA.
x <- read.tucson(tf)
x["1915", ]

## Ask for the old behaviour explicitly.
x0 <- read.tucson(tf, fill.internal.NA = 0)
x0["1915", ]

unlink(tf)

Run the code above in your browser using DataLab