Learn R Programming

rurl (version 3.0.1)

rurl-package: rurl: Parse, Clean, and Normalize URLs

Description

A lightweight toolkit for extracting structured information from URLs. Includes functions for parsing, normalizing protocols, extracting domains, and constructing clean URLs. Domain and public-suffix extraction is delegated to the 'pslr' package, which implements the Public Suffix List from https://publicsuffix.org. Punycode and IDNA encoding is handled by the 'punycoder' package.

Arguments

Author

Maintainer: Bart Turczynski [email protected] (ORCID)

Authors:

See Also

Parsing: safe_parse_url(), safe_parse_urls(). Accessors: get_host(), get_domain(), get_tld(), get_subdomain(), get_path(), get_query(). Cleaning and joining: get_clean_url(), canonical_join(). Query introspection: query_param_summary(). Cache management: rurl_clear_caches(), rurl_cache_info(), rurl_cache_config().

Domain and public-suffix extraction is delegated to the pslr package; Punycode/IDNA encoding is handled by the punycoder package.

Examples

Run this code
# Parse a vector of URLs into one row each.
urls <- c(
  "https://www.Example.co.uk/Blog/index.html?utm_source=nl&id=7#top",
  "http://sub.example.com:8080/a/./b/../c"
)
safe_parse_urls(urls)[, c("scheme", "host", "domain", "tld", "path")]

# Reach a single component without materializing the frame.
get_domain(urls)
get_subdomain(urls)

# Clean for SEO: a lossy projection of a WHATWG parse. Dot segments resolve,
# the host renders in Unicode, www/index come off, the query is dropped.
get_clean_url(urls, profile = "seo")

# Profiles are inspectable sugar over the low-level knobs, and an explicit
# argument always overrides the bundle.
url_profile("seo")
get_clean_url("https://xn--mnchen-3ya.de/a", profile = "seo")
get_clean_url("https://xn--mnchen-3ya.de/a", profile = "seo",
  host_encoding = "keep")

Run the code above in your browser using DataLab