Learn R Programming

rurl (version 3.0.1)

get_url_key: URL comparison key

Description

Projects each URL onto a versioned, non-URL comparison key: the value rurl uses to decide whether two URLs identify the same web resource. Use it to deduplicate, group, or match URLs without relying on a cleaned display string.

Usage

get_url_key(url, policy = url_key_policy())

Value

A classed character vector (rurl_url_key) the same length as url, preserving its names. NA for a non-keyable element. The policy version, schema version, standard, scheme-equality mode and the per-element keyability reasons ride along as attributes.

Arguments

url

A character vector of URLs. Factors are coerced.

policy

A rurl_url_key_policy object from url_key_policy(), which is also the default. The policy is scalar and is never recycled.

The key is not a URL

The returned object is a classed character vector whose contents are injectively framed component bytes, not a URL. Never parse it, never render it to users, and never reconstruct a URL from it. print() deliberately shows a truncated diagnostic form for that reason. What it is good for is comparison: ==, match(), duplicated(), %in% and the url_join family all work on it directly.

Framing is length-prefixed, so component boundaries cannot be forged. A host or path containing separators, control bytes or colons can never make two different URLs collide.

Non-keyable input

A URL the selected standard cannot parse has no identity, so its key is NA and never matches anything -- not even another NA. The reason is kept alongside rather than collapsed into the NA, and is readable with attr(key, "keyability"): one of "ok", "missing-input", "empty-input" or "invalid-parse". Missing input is never conflated with an invalid parse.

See Also

url_key_policy() for the dials, url_join for joining on the key, and serialize_url() for a standard's full-string serialization (which is a URL, unlike this).

Examples

Run this code
# Presentation differences that are not identity differences.
get_url_key(c("http://example.com:80/a", "http://example.com/a"))

# The fragment and userinfo are excluded from web-resource identity.
k <- get_url_key(c("http://u:[email protected]/a#top", "http://example.com/a"))
k[1] == k[2]

# Query order and duplicates are significant.
k <- get_url_key(c("http://example.com/?a=1&b=2",
                   "http://example.com/?b=2&a=1"))
k[1] == k[2]

# Deduplicate by identity rather than by string.
u <- c("HTTP://Example.com/a", "http://example.com/a",
       "http://example.com/b")
u[!duplicated(get_url_key(u))]

# Non-keyable input carries a typed reason.
k <- get_url_key(c("http://example.com/", NA, "", ":::"))
attr(k, "keyability")

Run the code above in your browser using DataLab