Learn R Programming

dplR (version 1.8.0)

write.tucson: Write Tucson Format Chronology File

Description

This function writes a chronology to a Tucson (decadal) format file.

Usage

write.tucson(rwl.df, fname, header = NULL, append = FALSE,
             prec = 0.01, mapping.fname = "", mapping.append = FALSE,
             long.names = FALSE, fill.internal.NA = NULL,
             extra.chars = c("-", "_", "."), ...)

Value

fname

Arguments

rwl.df

a data.frame containing tree-ring ring widths with the series in columns and the years as rows. The series IDs are the column names and the years are the row names. This type of data.frame is produced by read.rwl and read.compact.

fname

a character vector giving the file name of the rwl file.

header

a list giving information for the header of the file. If NULL then no header information will be written.

append

logical flag indicating whether to append this chronology to an existing file. The default is to create a new file.

prec

numeric indicating the precision of the output file. This must be equal to either 0.01 or 0.001 (units are in mm).

mapping.fname

a character vector of length one giving the file name of an optional output file showing the mapping between input and output series IDs. The mapping is only printed for those IDs that are transformed. An empty name (the default) disables output.

mapping.append

logical flag indicating whether to append the description of the altered series IDs to an existing file. The default is to create a new file.

long.names

logical flag indicating whether to allow long series IDs (7 or 8 characters) to be written to the output. The default is to only allow 6 characters.

fill.internal.NA

what to do with interior NA, i.e. years between a series' first and last measurement where rwl.df holds no value. The default, NULL, invents nothing and writes them as the missing-data sentinel -999 (see Details). Anything else is passed to fill.internal.NA before the file is written, so fill.internal.NA = 0 fills the gaps with zero and "Mean", "Spline" or "Linear" interpolate. This is the same argument, with the same meaning and default, as in read.tucson.

extra.chars

a character vector of characters to allow in series IDs in addition to a--z, A--Z and 0--9. The default allows the hyphen, the underscore and the period, all of which occur routinely in ITRDB IDs. Given either as single characters or as one string, so c("-", "_") and "-_" are the same request. Use extra.chars = character(0) to allow nothing beyond alphanumerics, which is what dplR did before version 1.8.0. Whitespace, control characters and "#" are refused: the first two break the fixed-width layout, and "#" is the default comment.char of read.tucson, which discards any line containing it.

...

Unknown arguments are accepted but not used.

Author

Andy Bunn. Patched and improved by Mikko Korpela.

Details

This writes a standard rwl file as defined according to the standards of the ITRDB at https://www.ncei.noaa.gov/pub/data/paleo/treering/treeinfo.txt. This is the decadal or Tucson format. It is an ASCII file and machine readable by the standard dendrochronology programs. Header information for the rwl can be written according to the International Tree Ring Data Bank (ITRDB) standard. The header standard is not very reliable however and should be thought of as experimental here. Do not try to write headers using dplR to submit to the ITRDB. When submitting to the ITRDB, you can enter the metadata via their website. If you insist however, the header information is given as a list and must be formatted with the following:

DescriptionNameClassMax Width
Site IDsite.idcharacter5
Site Namesite.namecharacter52
Species Codespp.codecharacter4
State or Countrystate.countrycharacter13
Speciessppcharacter18
Elevationelevcharacter5
Latitudelatcharacter or numeric5
Longitudelongcharacter or numeric5
First Yearfirst.yrcharacter or numeric4
Last Yearlast.yrcharacter or numeric4
Lead Investigatorlead.invscharacter63
Completion Datecomp.datecharacter8

See examples for a correctly formatted header list. If the width of the fields is less than the max width, then the fields will be padded to the right length when written. Note that lat and long are really lat * 100 or long * 100 and given as integral values. E.g., 37 degrees 30 minutes would be given as 3750.

Series can be appended to the bottom of an existing file with a second call to write.tucson. The output from this file is suitable for publication on the ITRDB.

The function is capable of altering excessively long and/or duplicate series IDs to fit the Tucson specification. Additionally, characters outside a--z, A--Z, 0--9 and extra.chars will be removed. If series IDs are changed, a single warning is shown giving the reasons and the first few renamings as old -> new. The user may also wish to print the full list (see mapping.fname in Arguments).

Renaming is worth attention because it is not confined to the series that prompted it. Removing a character shortens a name, and a shortened name can collide with one that needed no change; the duplicate pass then renames both. Before dplR 1.8.0, a file holding "CC1-1" and "CC11" was written out as "CC110" and "CC111", neither of which is an ID in the input.

The Tucson format does not restrict the character set of a series ID: the ITRDB description names only the columns the ID occupies. Hyphens, underscores and periods are common in real IDs, they do not disturb the fixed-width layout, and read.tucson reads them back unchanged, so they are written as they stand rather than removed. There is one position where a hyphen cannot go: column 8, which is where a minus sign distinguishes a 5-character year before -999 from an 8-character series ID. An 8-character ID ending in a hyphen is therefore shortened by that hyphen, with a warning. This can only arise when long.names = TRUE; at the default width an ID stops at column 6.

Interior NA -- years between a series' first and last measurement where rwl.df records nothing -- are written as -999, the negative value the Tucson format uses to mark missing data, at both precisions. read.tucson reads any negative value that is not the stop marker back as NA, so a read.tucson--write.tucson--read.tucson cycle preserves the gaps. Note that read.tucson.legacy, and other dendrochronology programs, will read those cells as a ring width of zero. A message is given whenever gaps are written.

Before dplR 1.8.0 interior NA were written as -999 at prec = 0.01 but as 0 at prec = 0.001. A zero ring width is a locally absent ring, which is a real observation about a tree in a year, and it is not the same statement as "not measured"; the two are no longer conflated. Use fill.internal.NA = 0 to get the old prec = 0.001 output. Files with no interior NA are written exactly as before.

A series that is entirely NA has nothing to write. It is skipped with a warning naming the series, and the rest of the file is written normally.

Setting long.names = TRUE allows series IDs to be 8 characters long, or 7 in case there are year numbers using 5 characters. Note that in the latter case the limit of 7 characters applies to all IDs, not just the one corresponding to the series with long year numbers. The default (long.names = FALSE) is to allow 6 characters. Long IDs may cause incompatibility with other software.

See Also

write.crn, read.tucson, read.tucson.legacy, fill.internal.NA, write.rwl, write.compact, write.tridas

Examples

Run this code
library(utils)
data(co021)
co021.hdr <- list(site.id = "CO021",
                  site.name = "SCHULMAN OLD TREE NO. 1, MESA VERDE",
                  spp.code = "PSME", state.country = "COLORADO",
                  spp = "DOUGLAS FIR", elev = "2103M", lat = 3712,
                  long = -10830, first.yr = 1400, last.yr = 1963,
                  lead.invs = "E. SCHULMAN", comp.date = "")
fname <- write.tucson(rwl.df = co021, fname = tempfile(fileext=".rwl"),
                      header = co021.hdr, append = FALSE, prec = 0.001)
print(fname) # tempfile used for output

unlink(fname) # remove the file

## Interior NA survive a write and a read
gappy <- data.frame(SER01 = round(seq(0.5, 2, length.out = 40), 3),
                    row.names = 1901:1940)
gappy[15:19, 1] <- NA  # five years with no measurement
fname2 <- write.tucson(gappy, tempfile(fileext=".rwl"), prec = 0.001)
back <- read.tucson(fname2, verbose = FALSE)
identical(is.na(back$SER01), is.na(gappy$SER01)) # TRUE

## The old behaviour, if you want it: gaps filled with zero
fname3 <- write.tucson(gappy, tempfile(fileext=".rwl"), prec = 0.001,
                       fill.internal.NA = 0)
read.tucson(fname3, verbose = FALSE)$SER01[15:19] # zeros, not NA

unlink(c(fname2, fname3))

## Hyphens and underscores in series IDs survive a round trip
dashed <- data.frame(`CC1-1` = 1:5, CC11 = 6:10, `CC_5` = 11:15,
                     `CC.7` = 16:20,
                     check.names = FALSE, row.names = 1901:1905)
fname4 <- write.tucson(dashed, tempfile(fileext=".rwl"))
names(read.tucson(fname4, verbose = FALSE))

## The pre-1.8.0 rule, if you want it: note that stripping the hyphen from
## CC1-1 collides with CC11, so both series are renamed
fname5 <- write.tucson(dashed, tempfile(fileext=".rwl"),
                       extra.chars = character(0))
names(read.tucson(fname5, verbose = FALSE))

unlink(c(fname4, fname5))

Run the code above in your browser using DataLab