This function serves as the core URL processing engine. It parses a URL, handles protocol and www prefix modifications, detects IP addresses, and derives components like the registered domain and top-level domain (TLD). Results are memoized for performance when processing large datasets.
safe_parse_url(
url,
protocol_handling = c("keep", "none", "strip", "http", "https"),
www_handling = c("none", "strip", "keep", "if_no_subdomain"),
tld_source = c("all", "private", "icann"),
case_handling = c("lower_host", "keep", "lower", "upper"),
trailing_slash_handling = c("none", "keep", "strip"),
index_page_handling = c("keep", "strip"),
path_normalization = c("none", "collapse_slashes", "dot_segments", "both"),
scheme_relative_handling = c("keep", "http", "https", "error"),
subdomain_levels_to_keep = NULL,
host_encoding = c("keep", "idna", "unicode"),
path_encoding = c("keep", "encode", "decode"),
query_handling = c("drop", "filter", "allow", "keep"),
params_keep = NULL,
params_drop = NULL,
sort_params = FALSE,
empty_param_handling = c("keep", "drop"),
params_case_sensitive = FALSE,
decode_plus = FALSE,
port_handling = c("exclude", "keep", "strip_default", "strip_all"),
scheme_policy = c("infer", "require"),
scheme_acceptance = c("web", "general"),
url_standard = NULL,
engine = NULL,
profile = NULL,
credential_handling = c("strip", "reject")
)A named list with the following components:
original_url: The original URL string provided.
scheme: The scheme (e.g., "http", "https").
host: The host (e.g., "www.example.com"). NA if the host becomes
empty after processing.
port: The port number.
path: The path component (e.g., "/path/to/resource").
query: The query string (e.g., "name=value"); never
percent-decoded. Under url_standard = "whatwg" it carries the
standard's percent-encoded spelling (the query percent-encode set is
applied, so a literal space reports as "%20"); under
url_standard = "rfc3986" or no selector it is the raw source spelling,
preserved byte-for-byte exactly as written in the URL (a bare key such
as "flag" stays "flag", not "flag="). A present-but-empty query (e.g.
from a trailing "?") is reported as NA.
fragment: The fragment identifier (e.g., "section"); never
percent-decoded, with the same two-branch contract as query (the
fragment percent-encode set is applied under url_standard = "whatwg",
so a double-quote inside the fragment reports as "%22"). Empty is
reported as NA.
user: The user name for authentication; never percent-decoded.
Under url_standard = "whatwg" it carries the standard's percent-encoded
spelling (the userinfo percent-encode set is applied, so
"http://a^b@host/" reports "a%5Eb"); under url_standard = "rfc3986" or
no selector it is the raw source spelling, exactly as written in the URL.
Empty is reported as NA.
password: The password for authentication, with the same
encoding contract as user (so a ":" inside a WHATWG password is
reported as "%3A"). Empty is reported as NA.
domain: The registered domain name (e.g., "example.com"). NA if
host is an IP, empty, or derivation fails.
tld: The top-level domain (e.g., "com"). NA if host is an IP,
empty, or derivation fails.
domain_ascii, domain_unicode: The registered domain in both
canonical spellings, independent of host_encoding. For an
internationalized domain, domain_ascii is the Punycode/A-label form
(e.g., "xn--mnchen-3ya.de") and domain_unicode the decoded Unicode form
(e.g., "münchen.de"); for an ASCII-only domain the two are equal.
Unlike domain (which follows host_encoding, a rendering choice),
these are stable identity keys — a Unicode host and its A-label share
one domain_ascii — so consumers can build an encoding-independent key
from a single parse. NA under the same conditions as domain.
tld_ascii, tld_unicode: The public suffix (TLD) in both
canonical spellings, the tld analogue of domain_ascii/
domain_unicode. NA under the same conditions as tld.
is_ip_host: Logical, TRUE if the host is an IP address.
clean_url: A normalized canonical key reconstructed from
scheme, host, and path, after processing and with case handling
applied. The query is included only when query_handling != "drop"
(the default is "drop", so by default the query is excluded); when
included it is filtered/canonicalized per the query options and appended
case-unfolded. The port is included only when
port_handling != "exclude" (the default is "exclude", so by default
the port is excluded, as before); fragment and userinfo are always
excluded (use the dedicated components above to retrieve them). With
path_encoding = "decode" the path is shown decoded, so clean_url
is human-readable rather than guaranteed URL-safe. NA on a parse error,
when the host is empty/NA except for a valid hostless file: URL, or
under credential_handling = "reject" when the authority carried a
userinfo delimiter.
parse_status: Character string indicating parsing outcome
("ok", "ok-ftp", "ok-scheme-relative", "error", "warning-no-tld",
"warning-invalid-tld", "warning-public-suffix", "warning-userinfo").
"warning-userinfo" marks a scheme-less input carrying userinfo (e.g.
"[email protected]"): host/domain/tld/user still resolve, but
clean_url is NA (rurl will not fabricate a canonical URL from an
ambiguous, email-shaped, scheme-less string).
Returns NULL if the URL is fundamentally unparseable (e.g., NA, empty)
or uses a disallowed scheme.
A single URL string to be parsed. For vectors, use
safe_parse_urls.
A character string specifying how to handle
protocols. Defaults to "keep".
Regardless of this option, rurl only processes authority-based URLs whose
scheme is one of http, https, ftp, or ftps; a scheme-bearing input with any
other scheme (e.g. mailto:, tel:, ws:) yields
parse_status = "error". Scheme inference (below) also requires the
input to be host-shaped: a scheme-less string that is not a host (e.g.
"asdfghjkl", "12345", "/path") or is a non-canonical
IP literal (integer/hex/octal/short forms, or leading-zero octets like
"192.168.010.1") is rejected as "error" rather than having a
scheme fabricated for it.
"keep": If a supported scheme exists (http, https, ftp, ftps), it's
used. If no scheme and the input is host-shaped, "http://" is added;
otherwise the input is not a URL and yields "error".
"none": If a supported scheme exists, it's used. If no scheme, then no scheme is used (scheme component will be NA).
"strip": Any existing scheme is removed (scheme component will be NA).
"http": The scheme is forced to be "http".
"https": The scheme is forced to be "https".
A character string specifying how to handle "www"
and www[number] prefixes in the host. Defaults to "none".
"none": (Default) Leaves the host's www prefix (or lack thereof) untouched.
"strip": Removes any "www." or www[number]. prefix.
"keep": Ensures the host starts with "www.". If it has
www[number]., it's normalized to "www.". If no www prefix, "www." is
added. An empty input host remains empty.
"if_no_subdomain": If the host is a bare registered domain (e.g.,
"example.com"), "www." is added. If the host already has a "www." or
www[number]. prefix, it is normalized to "www." (e.g.,
"www1.example.com" becomes "www.example.com"; "www1.sub.example.com"
becomes "www.sub.example.com"). If a non-www subdomain exists (e.g.,
"sub.example.com" or the normalized "www.sub.example.com"), the host is
not further altered. An empty input host remains empty.
Which TLD source to use for TLD extraction: "all", "icann", or "private". Defaults to "all".
A character string specifying how to handle the case of the cleaned URL. Defaults to "lower_host", the RFC 3986 §6.2.2.1 normalization (scheme and host are case-insensitive and folded to lowercase; the path is case-sensitive and preserved).
"lower_host": (Default) Lowercases scheme and host only; the path keeps its original casing.
"keep": Preserves casing of the reconstructed URL.
"lower": Converts the cleaned URL to lowercase.
"upper": Converts the cleaned URL to uppercase.
A character string specifying how to handle trailing slashes in the path component of the cleaned URL. Defaults to "none".
"none": (Default) No specific handling is applied. Path remains as is after initial parsing.
"keep": Ensures a trailing slash. If a path exists and doesn't end with one, it's added. If path is just "/", it's kept.
"strip": Removes a trailing slash if present, unless the path is solely "/".
A character string specifying how to handle index/default pages. Defaults to "keep".
"keep": (Default) Leave index/default page segments untouched.
"strip": Remove a trailing index.* or default.* segment (case-insensitive).
How to normalize path structure. Defaults to
"none". rurl owns dot-segment resolution: the path is read from the input
verbatim (never from a pre-normalized path), so "none" preserves
. / .. segments (/a/../b stays /a/../b) and only
the settings below change them. Resolution follows RFC 3986 section 5.2.4 and
acts on literal ./.. segments only — a percent-encoded
%2e is a normal path byte, never a dot segment, so it is never
treated as traversal.
"none": (Default) No normalization; dot and slash structure is preserved exactly as written.
"collapse_slashes": Collapse duplicate slashes in the path.
"dot_segments": Resolve . and .. segments per RFC 3986.
"both": Apply both collapse_slashes and dot_segments.
How to handle URLs starting with "//". Defaults to "keep".
"keep": Parse using http but return scheme as NA and set status to "ok-scheme-relative".
"http": Assume http for parsing and output.
"https": Assume https for parsing and output.
"error": Treat scheme-relative URLs as invalid.
An integer or NULL. Determines how many
levels of subdomains are kept,
in addition to any 'www.' prefix handled by www_handling.
NULL: (Default) No specific subdomain stripping is performed
beyond www_handling.
0: All subdomains are stripped. If www_handling preserved or
added 'www.',
it remains (e.g., 'www.sub.example.com' becomes 'www.example.com';
'sub.example.com' becomes 'example.com').
N > 0: Keeps up to N levels of subdomains, counted from
right-to-left (closest to the registered domain),
in addition to any 'www.' prefix. E.g., if N=1,
'three.two.one.example.com' becomes 'one.example.com';
'www.three.two.one.example.com' (post www_handling) becomes
'www.one.example.com'.
How to present the host in clean_url. Defaults to
"keep".
"keep": Leave the host as parsed (may preserve original case).
"idna": Convert Unicode host labels to Punycode (IDNA) for the cleaned URL.
"unicode": Decode Punycode labels to Unicode for the cleaned URL.
Under url_standard = "whatwg" every value renders the UTS-46-mapped
host, because mapping is part of WHATWG host parsing rather than a
feature of the idna dial (BÜCHER.example presents as
bücher.example; RUL-002). There "keep" preserves only whether the
input was written as an A-label (xn--...), so get_host() and
get_domain() agree on the same row. "rfc3986" and NULL are
unaffected.
How to present the path percent-encoding in clean_url
— the readable-vs-browser rendering choice (the path analog of
host_encoding). Defaults to "keep". This is an orthogonal presentation
knob: it is independent of url_standard and layers on top of any profile
(e.g. url_standard = "whatwg", path_encoding = "encode" emits the
WHATWG-parsed path in browser form), exactly like host_encoding. Only
"keep" preserves a profile's canonical identity path verbatim; "encode" and
"decode" are presentation forms that may re-encode or decode reserved octets
(so %2F may fold to a path-separating /), independent of whether a
profile is set.
"keep": Leave the path percent-encoding untouched (the path is
preserved as written in the URL, so %2F stays %2F rather than
decoding into a path-separating /). With no url_standard, rurl keeps
its historical RFC-style percent-hex case canonicalization, so %2f
becomes %2F. Under url_standard = "rfc3986", the profile's RFC 3986
§6.2.2.2 normalization applies: a triplet encoding an unreserved byte is
decoded, every other triplet stays encoded with uppercased hex, so
%7E becomes ~ while %2F stays %2F. Under
url_standard = "whatwg", existing percent-triplet
spelling is preserved byte-for-byte. Use "encode" to additionally
normalize which bytes are encoded.
"encode": The browser/percent-encoded rendering. Decodes the path first, then percent-encodes each segment (slashes preserved), so a readable non-ASCII path is emitted in its percent-encoded UTF-8 form.
"decode": The readable rendering. Percent-decodes UTF-8 sequences in the path, so a percent-encoded segment is shown as readable text.
A character string controlling whether (and how) the
query string is included in clean_url. Defaults to "drop", which preserves
the historical query-free clean_url. The raw query result field is never
affected by this option — it always reports the faithful original query.
"drop": (Default) clean_url carries no query, exactly as before.
"filter": Keep contentful params, dropping known trackers via a
built-in denylist (e.g. utm_*, fbclid, gclid). params_drop
extends the denylist; params_keep rescues names (winning over both the
denylist and empty-dropping).
"allow": Keep only params whose names match params_keep; all
others are dropped. Here params_keep is an inclusion criterion only,
not an empty-rescue.
"keep": Keep every param, re-encoded into canonical form (not the
verbatim original — that stays on the query field).
In every non-"drop" mode the surviving query is re-encoded canonically
(uppercase percent-hex, spaces as %20) and appended after the path. The
query is intentionally EXEMPT from case_handling (query values are
case-sensitive — tokens, IDs, signatures), so under
case_handling = "lower" or "upper" the clean_url is no longer
uniformly cased: scheme/host/path fold but the query keeps its original
case. Because clean_url is the canonical_join key, any
non-"drop" mode also brings the query into that join key (so ?id=1 and
?id=2 stop collapsing, while utm-only differences still collapse under
"filter").
Character vector of parameter-name globs (only * is
special), or NULL (default). In "filter" mode this is the rescue list; in
"allow" mode it is the allowlist. Ignored in "drop"/"keep".
Character vector of parameter-name globs to add to the
built-in denylist in "filter" mode, or NULL (default). Ignored in
"drop"/"allow"/"keep".
Logical (default FALSE). When TRUE, surviving params
are stably sorted by decoded key. Active in "filter"/"allow"/"keep".
One of "keep" (default) or "drop". "drop" removes
empty-valued params (e.g. ?ref=), except those rescued by params_keep
in "filter" mode.
Logical (default FALSE). Controls whether the
denylist and params_keep/params_drop matching is case-sensitive.
Logical (default FALSE). When TRUE, + in query
values is treated as a space (HTML-form decoding) before percent-decoding.
FALSE keeps + literal (RFC 3986 generic behavior).
A character string controlling whether the port
appears in clean_url. Defaults to "exclude", today's only historical
behavior. This knob is standalone and standard-independent (editorial, like
www_handling) -- url_standard never governs whether it may be set.
"exclude": (Default) The port never appears in clean_url.
"strip_all": Explicit alias of "exclude".
"keep": Include the syntactic port when present, including a
default port under url_standard = "whatwg". This is an explicit
non-parity override for callers that need the input's port spelling.
"strip_default": Keep only non-default ports (using the same
scheme-default table), independent of url_standard. Default-ness is
judged on the scheme the input was parsed with, never on the scheme
protocol_handling renders: http://example.com:443/a under
protocol_handling = "https" keeps :443, and http://example.com:80/a
drops :80 (RFC 3986 §6.2.3; WHATWG URL Standard port state; RUL-016).
This is the value profile = "seo" pins.
Controls whether scheme-less, host-shaped input is
accepted (an input-acceptance axis, distinct from protocol_handling,
which only controls how the scheme is presented, and from url_standard,
which controls interpretation). Defaults to "infer".
"infer": (Default) Fabricate http:// for scheme-less host-shaped
input (e.g. example.com parses as http://example.com), a
browser-omnibox-style affordance. This is the historical behavior.
"require": Reject scheme-less input — a scheme-less host-shaped
value becomes parse_status = "error" rather than gaining a fabricated
scheme. Use this for a strict, pure-parser posture. Note this governs
only bare host input; scheme-relative //host input is governed
separately by scheme_relative_handling.
Which scheme tokens may enter parsing (a
scheme-acceptance axis, distinct from scheme_policy, which governs
scheme-less input, and from url_standard, which governs interpretation).
Defaults to "web".
"web": (Default) Only the curated web-scheme allowlist
(http/https/ftp/ftps/file) is admitted; a scheme-bearing input
outside it is parse_status = "error". This is the historical,
byte-for-byte compatible behavior.
"general": Admit any syntactically valid scheme token and parse
opaque (mailto:x), non-special (foo://host), and RFC-generic URLs.
Requires an explicit url_standard ("rfc3986" or "whatwg"), which
decides the interpretation; general with url_standard = NULL is an
error. Non-special / opaque hosts receive no www-stripping, no domain/TLD
derivation, and are never run through the IDNA/punycode helpers.
A non-special scheme with no // is an opaque path: it has no
authority, so host, user, port and the domain/tld columns are
all NA and the entire remainder is the path (query/fragment are
still split off). This includes mailto: — the recipient's @ never
re-triggers authority parsing. To decompose a mailto: recipient, use
the accessors (get_host() / get_domain() / get_user(), ADR 0012 D7)
or get_mailto_recipients(); those deliberately return a recipient's
parts where this table presents NA, because a recipient domain is
extraction metadata, not the URL's authority.
Note that file: is admitted under both values, including the default.
A file: URL denotes local-filesystem access, and one with a non-empty
host (file://server/share/x) is a UNC path on Windows, so dereferencing
it reaches a remote SMB share. rurl parses file: URLs; it never opens
them. Restricting schemes before anything dereferences them is the
caller's job — see SECURITY.md.
Optional top-level standard profile: NULL (default),
"rfc3986", or "whatwg". With NULL the behavior is exactly what the
individual low-level options select (fully backward compatible). When set,
it selects a coherent set of standard-conformant behaviors for the axes it
governs — path percent/dot handling, the host IPv4/reg-name model, and
case_handling — so callers do not have to hand-assemble the low-level
knobs. Passing a governed low-level knob (path_normalization or
case_handling) with a value the selected
profile would not choose is an error; passing the value the profile would
pick is accepted (only case_handling = "lower_host" is accepted under a
selector — "keep", "lower", and "upper" all conflict, since "lower"
also lowercases the path, which neither standard sanctions). Added as the
last argument so existing positional calls keep their meaning; always
pass it by name. Under "whatwg" the selector additionally recognizes a
literal backslash as a path separator for WHATWG-special schemes
(http/https/ftp) and nulls default ports in parse output; use
port_handling = "strip_default" for spec-style clean URL port rendering.
See resolve_url for url_standard-governed
reference resolution. The selector does not govern whether
port_handling may be set (it is a standalone editorial knob), nor does it
govern path_encoding (an orthogonal path-presentation knob that layers
on any profile), IDNA rendering, or query handling.
Optional pslr engine controlling which Public Suffix List
backs domain / TLD / subdomain extraction: NULL (default) resolves
against pslr's session-global default list — exactly the historical
behavior — while a pslr::psl_engine() snapshot resolves against that
specific list, per request, without mutating any global state (never call
pslr::psl_use() for this). Use it to pin a particular list version or to
load an alternate list via pslr::psl_engine(source = "path", path = ...).
Process-local: an engine holds a C++ external pointer that does not
serialize across R sessions or parallel workers — build it in the process
that uses it; never cache it to disk or send it to a worker (rebuild one
per process instead). Only the domain-derived outputs (domain, tld,
and the subdomain-trimmed host / clean_url) depend on it.
Optional named profile bundling several knobs at once: NULL
(default; behaves exactly as the individual arguments select, fully
backward compatible), "browser", "whatwg", "rfc-syntax", "seo", or
the "seo" alias "canonical". A profile is separate from
url_standard (it bundles acceptance, interpretation, leniency, and
canonicalization together) and expands only into arguments you did not
supply explicitly — an explicit argument always overrides the profile.
"browser" is a browser-like fix-up posture (http-prepending; not
Chrome-faithful); "whatwg" is the absolute-URL no-base posture that
rejects scheme-less input (unlike a bare url_standard = "whatwg");
"rfc-syntax" is RFC 3986 generic syntax as parsing, not normalization
(case and dot-segments are preserved); "seo"/"canonical" is rurl's
origin-cleaning intent — a lossy policy projection of a WHATWG-parsed
URL (ADR 0017), which claims no resource equivalence:
url_standard = "whatwg" underneath (which also resolves ./.. folder
segments), https, a Unicode host regardless of the input spelling,
strip www / trailing slash / index page, drop the whole query, and drop
a default port only (port_handling = "strip_default": a
non-default port names a different origin and survives). Inspect
the resolved bundle with
url_profile. Also accepted by canonical_join
(forwarded through its ...).
How clean_url treats a URL whose parsed
authority carried a userinfo delimiter (user@, user:password@, a bare
@, or a repeated @). Defaults to "strip". A policy dial on the clean
surface (ADR 0017, mutation-table row 12; RUL-001), not a standards axis:
it composes with every url_standard, including NULL, and never
touches the user / password columns, parse_status, the diagnostics,
serialize_url or get_url_key.
"strip": (Default) The userinfo is dropped and the rest of the URL is emitted, exactly as before this argument existed.
"reject": clean_url is NA for such a row. RFC 3986 section
3.2.1 deprecates the user:password form and lets an application
reject it; sections 7.5 and 7.6 describe the credential leak and the
https://[email protected]/ semantic attack a silently
collapsed clean URL would hide. Use this when a cleaned URL that
looks like the credential-free original would be misleading.
There is no "keep": serialize_url already preserves
credentials under both standards, and format_url redacts
them for display.
safe_parse_urls
safe_parse_url(
"http://www.Example.com/Path?q=1#Frag",
protocol_handling = "keep",
case_handling = "lower"
)
safe_parse_url(
"Example.com/Another",
protocol_handling = "none",
www_handling = "keep",
case_handling = "upper",
trailing_slash_handling = "keep"
)
safe_parse_url(
"example.com",
www_handling = "if_no_subdomain"
) # -> www.example.com
safe_parse_url(
"sub.example.com",
www_handling = "if_no_subdomain"
) # -> sub.example.com
safe_parse_url(
"www1.example.com",
www_handling = "if_no_subdomain"
) # -> www.example.com
safe_parse_url(
"www1.sub.example.com",
www_handling = "if_no_subdomain"
) # -> www.sub.example.com
safe_parse_url(
"http://www.example.com/path/",
trailing_slash_handling = "strip"
)
safe_parse_url("192.168.1.1/test")
safe_parse_url("ftp://user:[email protected]:21/file.txt")
safe_parse_url(
"http://deep.sub.domain.example.com",
subdomain_levels_to_keep = 0
)
safe_parse_url(
"http://deep.sub.domain.example.com",
subdomain_levels_to_keep = 1
)
safe_parse_url(
"http://www.deep.sub.domain.example.com",
www_handling = "keep",
subdomain_levels_to_keep = 0
)
safe_parse_url(
"http://www.deep.sub.domain.example.com",
www_handling = "keep",
subdomain_levels_to_keep = 1
)
# Query handling: keep contentful params, drop known trackers.
safe_parse_url(
"http://example.com/watch?v=abc&utm_source=nl",
query_handling = "filter"
)$clean_url
# -> "http://example.com/watch?v=abc"
# params_keep is a RESCUE in "filter" (wins over the denylist) ...
safe_parse_url(
"http://example.com/?utm_source=nl&id=1",
query_handling = "filter", params_keep = "utm_source"
)$clean_url
# -> "http://example.com/?utm_source=nl&id=1"
# ... but an ALLOWLIST in "allow" (only listed names survive).
safe_parse_url(
"http://example.com/?a=1&id=2",
query_handling = "allow", params_keep = "id"
)$clean_url
# -> "http://example.com/?id=2"
# "allow" empty-handling asymmetry: params_keep does NOT rescue empties, so
# an allowed empty param still drops under empty_param_handling = "drop".
safe_parse_url(
"http://example.com/?id=&keep=1",
query_handling = "allow", params_keep = c("id", "keep"),
empty_param_handling = "drop"
)$clean_url
# -> "http://example.com/?keep=1"
Run the code above in your browser using DataLab