- data
A data frame or tibble. Occurrence records as input.
- path
A directory as character. The path to a folder that the output can be saved.
- taxonomy
A data frame or tibble. The taxonomy file to use.
Default = beesTaxonomy(); for other taxa see first taxadbToBeeBDC().
- speciesColumn
Character. The name of the column containing species names. Default = "scientificName".
- rm_names_clean
Logical. If TRUE then the names_clean column will be removed at the end of
this function to help reduce confusion about this column later. Default = TRUE
- checkVerbatim
Logical. If TRUE then the verbatimScientificName will be checked as well
for species matches. This matching will ONLY be done after harmoniseR has failed for the other
name columns. NOTE: this column is not first run through bdc::bdc_clean_names. Default = FALSE
- relaxAmbiguous
Logical. If TRUE, then ambiguous names will be returned as "TRUE" for
.invalidName. However, they will be listed as "ambiguouslyMatchedName"
under the "scientificNameAuthorship" column and the speciesColumn will NOT include the authority
information. Hence, the names can be used but in should remain clear in the data that they must
be used with caution. Default = FALSE.
- matchHigherTaxonomy
Logical. If TRUE, then BeeBDC will try to match higher taxonomies
(e.g., family names) using genera and such. This will work better if you have run the data through
bdc::bdc_clean_names beforehand. The function will update the taxonRank column, if it is
otherwise empty. Default = TRUE.
- stepSize
Numeric. The number of occurrences to process in each chunk. Default = 1000000.
- mc.cores
Numeric. If > 1, the function will run in parallel
using mclapply using the number of cores specified. If = 1 then it will be run using a serial
loop. NOTE: Windows machines must use a value of 1 (see ?parallel::mclapply). Additionally,
be aware that each thread can use large chunks of memory.
Default = 1.