Learn R Programming

rurl (version 3.0.1)

resolve_url: Resolve a URL reference against a base URL

Description

Resolves a relative or absolute URL reference against a base URL following the RFC 3986 section 5 reference-resolution algorithm, then renders the resolved absolute URL on the output surface output selects. Under the default output = "clean" the result is canonicalized with the same machinery as safe_parse_url; under output = "serialized" it is handed to serialize_url instead. url_standard and any ... options flow straight through to the parse, so the host IPv4/reg-name model, path percent/dot-segment handling, default-port elision, WHATWG backslash-as-slash recognition, and diagnostics are exactly those of a direct safe_parse_url() call on the resolved URL.

Usage

resolve_url(
  relative_or_absolute,
  base_url,
  url_standard = NULL,
  output = c("clean", "serialized"),
  form = c("source", "normalized"),
  ...
)

Value

A character vector the same length as the recycled inputs, unnamed (names are not data). Under output = "clean" each element is the canonical clean_url of the resolved reference; under output = "serialized" it is the standard's full-string serialization of the resolved absolute URL. NA where resolution cannot produce an absolute URL, or where the resolved URL is not accepted by the parser ("clean") or by the selected standard's parser ("serialized").

Arguments

relative_or_absolute

A character vector of URL references to resolve. Each may be relative ("../b", "?q=1", "#frag", "//host/p") or already absolute ("https://host/p"); an absolute reference ignores base_url.

base_url

A character vector of base URLs, recycled against relative_or_absolute. Each base must itself be an absolute URL (carry a scheme); a relative reference resolved against a scheme-less or NA base yields NA.

url_standard

Optional standard profile forwarded to the parse: NULL (default), "rfc3986", or "whatwg". See safe_parse_url for the axes it governs, and Reference resolution is standard-aware for the reference-parsing rules "whatwg" adds ahead of the merge. Required (non-NULL) when output = "serialized".

output

Which output surface to return: "clean" (default, today's canonical clean_url bytes) or "serialized" (the selected standard's full-string serialization of the resolved absolute URL, via serialize_url). See Which output surface you want.

form

For output = "serialized" with url_standard = "rfc3986" only, the RFC posture forwarded to serialize_url: "source" (default, source-preserving) or "normalized". Ignored for "whatwg", whose serializer has a single spec-defined form, and ignored under output = "clean", whose rendering is driven by the cleaning dials instead -- the same argument-is-inert-where-it-does-not-apply contract serialize_url itself holds for form.

...

Additional arguments forwarded to safe_parse_urls (e.g. port_handling, query_handling, host_encoding). Passing a governed low-level knob that conflicts with url_standard errors, exactly as it does for safe_parse_url. These are presentation dials consumed by the "clean" path only; supplying any of them together with output = "serialized" is an error, because serialize_url takes no presentation arguments and the dial could not be honored.

Reference resolution is standard-aware

The merge itself (empty reference, fragment-only, query-only, scheme-relative //host reference, absolute-path reference, and relative-path merge) is RFC 3986 section 5.2--5.3 under url_standard = "rfc3986" and under the default NULL selector, with one exception under "rfc3986" noted last below. Under url_standard = "whatwg" the WHATWG URL Standard's reference-parsing rules are applied first, because they are rules the two standards genuinely disagree on rather than composition of the axes url_standard already governs (decision P2.7 D-B, design/work/url-v3/decisions/P2.7-display-and-resolver-output.md):

  • A reference carrying the base's own special scheme is relative, not absolute. WHATWG's “special relative or authority state” consumes a scheme equal to the base's when that scheme is special (http, https, ws, wss, ftp, file) and keeps parsing against the base, so resolve_url("http:foo.com", "http://example.org/foo/bar", url_standard = "whatwg") is "http://example.org/foo/foo.com", while RFC 3986 treats any scheme as making the reference absolute and gives "http://foo.com/". A different scheme stays absolute under both, even when it is also special.

  • Under a special base, a leading \ in the reference is a /, and a run of either introduces an authority. WHATWG's “relative slash state” reads \ exactly as /, so resolve_url("\x", "http://example.org/foo/bar", url_standard = "whatwg", output = "serialized") is "http://example.org/x"; a second slash-or-backslash enters the authority, and “special authority ignore slashes” then skips the whole run for the five non-file special schemes, so "///example.org/x" and "/\/\//example.org/x" both resolve with the host example.org. file has its own state machine and consumes exactly two. RFC 3986 has neither rule: \ is an ordinary path byte and // is the entire authority production.

  • The reference is stripped before it is read. The WHATWG basic URL parser's step 1 removes a leading and trailing run of C0 control or SPACE (U+0000--U+0020) and then every ASCII tab, LF and CR, so " foo.com " resolves as "foo.com" does. A reference that strips to the empty string is the empty reference, which resolves to the base minus its fragment. RFC 3986 has no strip step -- such bytes are required to be percent-encoded -- so under "rfc3986" and NULL they stay in the reference.

  • Against a file: base, Windows drive letters follow the WHATWG file: state machine. A reference that begins with a drive letter empties the base path instead of shortening it, so resolve_url("C|/foo", "file:///tmp/mock/path", url_standard = "whatwg", output = "serialized") is "file:///C:/foo"; a rooted reference inherits the base's drive letter ("/" against "file:///C:/a/b" is "file:///C:/"); .. never removes a lone drive letter (".." against "file:///C:/" is "file:///C:/"); and C| in the first segment is normalized to C:. A drive letter in the authority position ("//d:") is an empty host plus a path segment, "file:///d:". RFC 3986 has no drive-letter concept, so under "rfc3986" and NULL the plain section 5.2 merge applies.

The NULL selector is frozen and unaffected (ADR 0007; P2.7 D-C): every rule above is reachable only through url_standard = "whatwg".

Two further rules apply under both named profiles, because the two standards agree on them. First, a resolved path whose first segment is empty is recomposed with the /. guard: RFC 3986 section 3.3 forbids a path beginning with // after no authority, and the WHATWG URL serializer emits the same guard, so resolve_url("/..//path", "non-spec:/p", url_standard = "whatwg", output = "serialized") is "non-spec:/.//path" rather than a string that re-reads as the authority path. The NULL selector recomposes the unguarded string, as it always did. Second, a scheme is ALPHA *( ALPHA / DIGIT / "+" / "-" / "." ) -- RFC 3986 section 3.1's own grammar, and WHATWG's -- so a relative path whose first segment merely contains a colon is a path, not an absolute reference: resolve_url("[61:24:74]:98", "http://example.org/foo/bar", url_standard = "whatwg", output = "serialized") is "http://example.org/foo/[61:24:74]:98", and "rfc3986" merges the same way. The NULL selector instead keeps RFC 3986 Appendix B's explicitly non-validating [^:/?#]+, which reads "10.0.0.7" as a scheme and discards the base; that is frozen behavior (ADR 0007), not a recommendation.

Which output surface you want

output selects between two different products, not two settings of one (decision P2.7 D-A, design/work/url-v3/decisions/P2.7-display-and-resolver-output.md):

  • output = "clean" (the default) returns the canonical clean_url of the resolved reference, not a verbatim RFC 3986 recomposition: as everywhere else in rurl, the fragment and userinfo are excluded from clean_url, the query is included only when query_handling != "drop" (the default drops it), and the port only when port_handling != "exclude". This surface is intentionally lossy -- it is a cleaning/SEO product driven by presentation policy, and it therefore cannot carry a conformance claim. This differs from a generic resolver such as xml2::url_absolute() or Python's urljoin, which preserve every component verbatim; resolve_url() resolves and canonicalizes.

  • output = "serialized" returns serialize_url(<resolved absolute URL>, standard = url_standard, form = form): the standard's own full-string serialization, with the fragment preserved and credentials reconstructed. This is the standards surface -- the one a conformance claim may be measured on -- and it is where RFC 3986 section 5.4's own expectations are reproduced exactly (resolve_url("?y", "http://a/b/c/d;p?q", url_standard = "rfc3986", output = "serialized") is "http://a/b/c/d;p?y").

output = "serialized" requires an explicit url_standard: NULL selects no standard, so there is nothing to serialize to, and the combination is an error rather than a silent choice of one. Because serialize_url accepts no presentation options at all, output = "serialized" also rejects any ... argument: honoring, say, port_handling = "exclude" is impossible on that surface, and accepting-then-discarding it would misreport what was returned.

To inspect individual resolved components (including the fragment), resolve first and pass the result to safe_parse_url.

See Also

safe_parse_url, get_clean_url, serialize_url

Examples

Run this code
resolve_url("../g", "http://a/b/c/d;p?q") # -> "http://a/b/g"
resolve_url("g", "http://a/b/c/d;p?q") # -> "http://a/b/c/g"
resolve_url("//example.org/p", "http://a/b/c") # -> "http://example.org/p"
resolve_url("https://x.com/y", "http://a/b/c") # absolute ref, base ignored
resolve_url(c("g", "../h"), "http://a/b/c/") # vectorized

# The standards surface keeps the query and the fragment RFC 3986 section
# 5.4 requires; the (lossy) clean surface drops both by design.
resolve_url("?y", "http://a/b/c/d;p?q",
            url_standard = "rfc3986", output = "serialized")
resolve_url("#s", "http://a/b/c/d;p?q",
            url_standard = "whatwg", output = "serialized")
resolve_url("#s", "http://a/b/c/d;p?q")

Run the code above in your browser using DataLab