Learn R Programming

DCC (version 1.2.1)

dcc_detect_encoding: Detect the character encoding of a text file

Description

Reads up to n_bytes from the file and detects the most likely character encoding via stringi::stri_enc_detect(). Detected encodings are normalized to the canonical names DCC supports as first-class: "UTF-8", "GB18030" (covers GBK/GB2312), "BIG5", and "latin1" (covers ISO-8859-1/windows-1252).

Usage

dcc_detect_encoding(path, n_bytes = 65536L)

Value

A list with elements encoding (normalized name), confidence (0-1), and candidates (data.frame of raw detector output).

Arguments

path

Path to the file.

n_bytes

Maximum number of bytes to sample (default 65536).

Examples

Run this code
f <- tempfile(fileext = ".csv")
writeLines("id,name", f)
dcc_detect_encoding(f)$encoding

Run the code above in your browser using DataLab