Learn R Programming

dplR (version 1.8.0)

read.sheet: Read Ring Widths from a Spreadsheet File

Description

Read a file of ring widths saved in “spreadsheet” layout: years down the rows, series across the columns, and the years in the first column.

Usage

read.sheet(fname, sep = NULL, dec = ".", layout = c("wide", "long"),
           transpose = FALSE, comment.char = "#", fill.internal.NA = NULL,
           fix.dup.char = "X", encoding = NULL, verbose = TRUE,
           strict = FALSE, ...)

csv2rwl(fname, ...)

Value

An object of class c("rwl", "data.frame") with the series in columns and the years as rows. The series IDs are the column names and the years are the row names.

The object carries an attribute "dplR.provenance", a list in the same shape as the one attached by read.tucson, recording what the reader found and what it changed. Its events element is a

data.frame of findings, each with a stable id such as

"ID_RENAMED" or "UNITS_SUSPECT".

Arguments

fname

a character vector giving the file name.

sep

the field separator: one of ",", "\t", ";" or "|", or any single character. NULL, the default, detects it.

dec

the decimal mark, "." or ",". A decimal comma is detected automatically when the separator is a semicolon.

layout

the shape of the file. "wide", the default, is years down the rows and series across the columns. "long" is a “long” or “tidy” file of series, year, value triples. Note that this is unrelated to the long argument of read.crn and read.tucson, which refers to a wider fixed-width year field.

transpose

logical flag. Set to TRUE for a file with the series in rows and the years in columns.

comment.char

lines beginning with this string are treated as a comment header: skipped, and kept in the provenance record.

fill.internal.NA

passed to fill.internal.NA to fill interior gaps. The default, NULL, leaves them as NA.

fix.dup.char

character string appended to a series ID that appears more than once in the header row.

encoding

the encoding of the file, used only if the file is not valid UTF-8. Any name iconv accepts will do, e.g. "latin1" or "ISO-8859-2". NULL, the default, means no encoding has been declared: the reader then falls back to "latin1" and says so. The tiers, and why the fallback does not run a charset detector, are described under ‘Encoding’ in read.tucson; read.sheet behaves identically.

verbose

logical flag. Write a summary of the file to the console.

strict

logical flag. If TRUE, every recoverable problem becomes an error instead of a warning.

...

other arguments passed to fread.

Deprecated

csv2rwl is deprecated as of dplR 1.8.0 and now calls read.sheet. It will be made defunct in a later release.

This is a change of behaviour and not only a change of name. csv2rwl assigned the class "rwl" directly, which bypassed as.rwl and its requirement that the row names be consecutive integers. It therefore accepted files that read.sheet refuses:

  • non-consecutive or duplicated years, which produced an object whose time was not a continuous span;

  • columns holding text, which travelled inside the rwl until some later function failed on them;

  • series IDs were passed through read.table(check.names = TRUE), so "1A" became "X1A" and "LF-2B" became "LF.2B", silently.

A file that csv2rwl read may therefore now stop with an error. That is the deprecation doing its job: the object it used to return was invalid. csv2rwl also stopped when a series ID appeared twice, where read.sheet renames the duplicate and records the change, as read.tucson does.

There is no csv2rwl.legacy. read.tucson.legacy exists because the behaviour of the old Tucson reader differs from the new one in ways that are occasionally wanted; that is not the case here.

Author

Andy Bunn

Details

The expected layout is the one a spreadsheet produces: the first column holds the years, the first row holds the series IDs, and each remaining cell holds one ring width.

YearSer1ASer1BSer2ASer2B
1901NA0.450.430.24
1902NA0.050.000.07
19030.170.460.030.21
19040.280.210.540.41

An empty cell, NA, NaN, "." and "-" are all read as missing. Any other value that does not parse as a number is an error rather than a missing value, and the offending cells are named.

Unlike a Tucson file, a spreadsheet has no format to conform to, so this function checks the contents rather than assuming them. The years must be whole numbers, ascending, unique and consecutive: a year with no measurements is a row of NA, not a missing row. Every column after the first must be numeric. Series IDs are read verbatim, and anything that changes one -- a stripped byte order mark, a resolved duplicate -- is reported and recorded rather than done silently.

Separators

Comma, tab, semicolon and pipe separated files are all read. Left at NULL, sep is detected by looking for the candidate that gives every line the same number of fields, preferring the one that gives the most; quoted spans are ignored while counting, so a series ID may contain the separator. The detected separator is recorded in the provenance record as SEP_GUESS and printed when verbose is TRUE, but it is not a warning: detection is the default, and a defect it is not.

A semicolon separated file is usually a European export and carries decimal commas. When the separator is detected as a semicolon and a digit-comma-digit appears in the data, dec is set to "," and recorded as DEC_COMMA. Thousands separators are not handled: "1.234,5" and "1,234.5" are the same characters under two conventions and nothing in the file says which.

sep and dec may not be the same character.

Long format

With layout = "long" the file holds one row per observation rather than a rectangle: a series ID, a year and a value. Columns are taken by position, which is what write.sheet writes; if the header names all three of series, year and value, in any case and any order, the names are used instead.

The year span of the result runs from the earliest to the latest year in the file. A year with no observation in any series therefore comes back as a row of NA, and is reported, since years with no data break several dplR functions. One series may hold only one value per year.

read.sheet refuses a NOAA/NCEI template file. Those files are a commented metadata block followed by a tab separated data table, and reading the table alone would discard the coordinates, species, investigators and DOI above it.

See Also

write.sheet, read.rwl, read.tucson, rwl.check, as.rwl

Examples

Run this code
library(utils)
data(ca533)
# write out a sheet that read.sheet will understand
tm <- time(ca533)
foo <- data.frame(tm, ca533, check.names = FALSE)
names(foo)[1] <- "Year"
tmpName <- tempfile(fileext = ".csv")
write.csv(foo, file = tmpName, row.names = FALSE)

# read it back in
bar <- read.sheet(tmpName, verbose = FALSE)
# the provenance record travels with the data
str(attr(bar, "dplR.provenance")$events)

unlink(tmpName)

Run the code above in your browser using DataLab