Read a file of ring widths saved in “spreadsheet” layout: years down the rows, series across the columns, and the years in the first column.
read.sheet(fname, sep = NULL, dec = ".", layout = c("wide", "long"),
transpose = FALSE, comment.char = "#", fill.internal.NA = NULL,
fix.dup.char = "X", encoding = NULL, verbose = TRUE,
strict = FALSE, ...)csv2rwl(fname, ...)
An object of class c("rwl", "data.frame") with the series in columns
and the years as rows. The series IDs are the column names and the
years are the row names.
The object carries an attribute "dplR.provenance", a list in the
same shape as the one attached by read.tucson, recording what
the reader found and what it changed. Its events element is a
data.frame of findings, each with a stable id such as
"ID_RENAMED" or "UNITS_SUSPECT".
a character vector giving the file name.
the field separator: one of ",", "\t",
";" or "|", or any single character. NULL, the
default, detects it.
the decimal mark, "." or ",". A decimal comma is
detected automatically when the separator is a semicolon.
the shape of the file. "wide", the default, is years
down the rows and series across the columns. "long" is a
“long” or “tidy” file of series, year,
value triples. Note that this is unrelated to the long
argument of read.crn and read.tucson, which
refers to a wider fixed-width year field.
logical flag. Set to TRUE for a file with the
series in rows and the years in columns.
lines beginning with this string are treated as a comment header: skipped, and kept in the provenance record.
passed to fill.internal.NA to fill
interior gaps. The default, NULL, leaves them as NA.
character string appended to a series ID
that appears more than once in the header row.
the encoding of the file, used only if the file is not
valid UTF-8. Any name iconv accepts will do,
e.g. "latin1" or "ISO-8859-2". NULL, the default,
means no encoding has been declared: the reader then falls back to
"latin1" and says so. The tiers, and why the fallback does not
run a charset detector, are described under ‘Encoding’ in
read.tucson; read.sheet behaves identically.
logical flag. Write a summary of the file to the
console.
logical flag. If TRUE, every recoverable problem
becomes an error instead of a warning.
other arguments passed to fread.
csv2rwl is deprecated as of dplR 1.8.0 and now calls
read.sheet. It will be made defunct in a later release.
This is a change of behaviour and not only a change of name. csv2rwl
assigned the class "rwl" directly, which bypassed as.rwl
and its requirement that the row names be consecutive integers. It therefore
accepted files that read.sheet refuses:
non-consecutive or duplicated years, which produced an object whose
time was not a continuous span;
columns holding text, which travelled inside the rwl until
some later function failed on them;
series IDs were passed through
read.table(check.names = TRUE), so "1A" became
"X1A" and "LF-2B" became "LF.2B", silently.
A file that csv2rwl read may therefore now stop with an error. That is
the deprecation doing its job: the object it used to return was invalid.
csv2rwl also stopped when a series ID appeared twice, where
read.sheet renames the duplicate and records the change, as
read.tucson does.
There is no csv2rwl.legacy. read.tucson.legacy exists
because the behaviour of the old Tucson reader differs from the new one in
ways that are occasionally wanted; that is not the case here.
Andy Bunn
The expected layout is the one a spreadsheet produces: the first column holds the years, the first row holds the series IDs, and each remaining cell holds one ring width.
| Year | Ser1A | Ser1B | Ser2A | Ser2B |
| 1901 | NA | 0.45 | 0.43 | 0.24 |
| 1902 | NA | 0.05 | 0.00 | 0.07 |
| 1903 | 0.17 | 0.46 | 0.03 | 0.21 |
| 1904 | 0.28 | 0.21 | 0.54 | 0.41 |
An empty cell, NA, NaN, "." and "-" are all read
as missing. Any other value that does not parse as a number is an error
rather than a missing value, and the offending cells are named.
Unlike a Tucson file, a spreadsheet has no format to conform to, so this
function checks the contents rather than assuming them. The years must be
whole numbers, ascending, unique and consecutive: a year with no measurements
is a row of NA, not a missing row. Every column after the first must
be numeric. Series IDs are read verbatim, and anything that changes
one -- a stripped byte order mark, a resolved duplicate -- is reported and
recorded rather than done silently.
Comma, tab, semicolon and pipe separated files are all read. Left at
NULL, sep is detected by looking for the candidate that
gives every line the same number of fields, preferring the one that gives
the most; quoted spans are ignored while counting, so a series
ID may contain the separator. The detected separator is recorded
in the provenance record as SEP_GUESS and printed when
verbose is TRUE, but it is not a warning: detection is the
default, and a defect it is not.
A semicolon separated file is usually a European export and carries decimal
commas. When the separator is detected as a semicolon and a digit-comma-digit
appears in the data, dec is set to "," and recorded as
DEC_COMMA. Thousands separators are not handled: "1.234,5"
and "1,234.5" are the same characters under two conventions and
nothing in the file says which.
sep and dec may not be the same character.
With layout = "long" the file holds one row per observation rather
than a rectangle: a series ID, a year and a value. Columns are taken
by position, which is what write.sheet writes; if the header
names all three of series, year and value, in any
case and any order, the names are used instead.
The year span of the result runs from the earliest to the latest year in
the file. A year with no observation in any series therefore comes back as
a row of NA, and is reported, since years with no data break several
dplR functions. One series may hold only one value per year.
read.sheet refuses a NOAA/NCEI template file. Those
files are a commented metadata block followed by a tab separated data table,
and reading the table alone would discard the coordinates, species,
investigators and DOI above it.
write.sheet, read.rwl,
read.tucson, rwl.check, as.rwl
library(utils)
data(ca533)
# write out a sheet that read.sheet will understand
tm <- time(ca533)
foo <- data.frame(tm, ca533, check.names = FALSE)
names(foo)[1] <- "Year"
tmpName <- tempfile(fileext = ".csv")
write.csv(foo, file = tmpName, row.names = FALSE)
# read it back in
bar <- read.sheet(tmpName, verbose = FALSE)
# the provenance record travels with the data
str(attr(bar, "dplR.provenance")$events)
unlink(tmpName)
Run the code above in your browser using DataLab