Learn R Programming

rurl (version 3.0.1)

url_join: Identity-keyed URL joins

Description

Six joins that match rows on URL identity rather than on string equality. Each side names one URL column; both sides are keyed with one immutable url_key_policy(), and rows pair up when their comparison keys are equal.

Usage

url_inner_join(
  x,
  y,
  by,
  policy = url_key_policy(),
  suffix = c(".x", ".y"),
  key_name = NULL,
  relationship = "none",
  multiple = "all",
  invalid = "keep",
  warnings = "allow",
  engine = NULL
)

url_left_join( x, y, by, policy = url_key_policy(), suffix = c(".x", ".y"), key_name = NULL, relationship = "none", multiple = "all", invalid = "keep", warnings = "allow", engine = NULL )

url_right_join( x, y, by, policy = url_key_policy(), suffix = c(".x", ".y"), key_name = NULL, relationship = "none", multiple = "all", invalid = "keep", warnings = "allow", engine = NULL )

url_full_join( x, y, by, policy = url_key_policy(), suffix = c(".x", ".y"), key_name = NULL, relationship = "none", multiple = "all", invalid = "keep", warnings = "allow", engine = NULL )

url_semi_join( x, y, by, policy = url_key_policy(), key_name = NULL, relationship = "none", invalid = "keep", warnings = "allow", engine = NULL )

url_anti_join( x, y, by, policy = url_key_policy(), key_name = NULL, relationship = "none", invalid = "keep", warnings = "allow", engine = NULL )

Value

A data frame built by row-slicing x, so x's column types and subclass survive. url_inner_join(), url_left_join(), url_right_join() and url_full_join() return x's columns followed by y's, disambiguated by suffix; url_semi_join() and url_anti_join()

return x's columns only. A zero-row result is built by the same path, so it carries the complete typed schema.

Arguments

x, y

Data frames to join.

by

The URL columns to key on: either one column name present on both sides ("URL"), or the named form c(x_col = "y_col") when they differ.

policy

A rurl_url_key_policy from url_key_policy(), applied symmetrically to both sides. Side-specific rules are prohibited: equality has to stay symmetric and transitive.

suffix

Length-2 character vector disambiguating column names present on both sides. Default c(".x", ".y"). If the result would still contain a duplicate name, the join errors rather than repairing it silently.

key_name

Optional column name under which to expose the comparison key. NULL (default) hides it. The exposed value is the classed key from get_url_key(), never a URL-looking string, and a name that collides with an output column is an error.

relationship

Cardinality you assert about matching keys, checked before the result is materialized: "none" (default, no check), "one-to-one", "one-to-many", "many-to-one", or "many-to-many" (no constraint, declared explicitly).

multiple

How many y rows a matching x row may take: "all" (default, lossless) or the separately named lossy narrowings "first" / "last", which take the first or last match in y row order.

invalid

What to do with rows that cannot be keyed: "keep" (default), "drop" or "error".

warnings

What to do with rows that parsed with a warning: "allow" (default), "reject" (ineligible to match, but retained) or "error".

engine

Optional psl_engine object from pslr::psl_engine() for per-request Public Suffix List resolution. NULL (default) uses the session-global engine. It cannot affect the comparison key -- identity frames no public-suffix component -- but it can affect which rows count as warning rows under warnings = "reject".

Why not a plain join

Joining data frames on raw URL strings misses http://example.com:80/a against http://example.com/a. Joining them on a cleaned string overmatches instead, because cleaning is a display policy: it can strip a trailing slash, a query parameter or a www. that genuinely distinguished two resources. These joins use get_url_key(), so what matches is what rurl considers the same resource -- and no cleaning or display option can change that.

Row order

Order is part of the contract, not an artifact of the implementation:

url_inner_join

matching pairs in x order, y matches in y order within each x row.

url_left_join

every x row in x order; unmatched x rows carry a missing y payload typed from y's own columns.

url_right_join

the exact mirror: every y row in y order, x matches in x order.

url_full_join

the left-join result, then the y rows it never consumed, in y order.

url_semi_join

each x row with at least one match, once, x columns only.

url_anti_join

each x row with no match, once, x columns only.

Duplicate keys expand as a Cartesian product. Rows are never silently discarded to "resolve" a duplicate -- multiplicity is a fact you declare with relationship or narrow with multiple.

Rows that cannot be keyed

A URL the standard cannot parse, or a missing or empty one, has no identity and never matches -- not even another unparseable URL. invalid decides what happens to those rows: "keep" (default) leaves them in, unmatched, so a left join still returns them; "drop" removes them before matching; "error" refuses the join and reports the offending row positions.

warnings is a separate axis for rows that did parse but carry a note -- userinfo on a scheme-less input, or a host whose public-suffix annotation did not resolve. "allow" (default) matches them normally, "reject" makes them ineligible to match without removing them, and "error" refuses the join.

url_anti_join() keeps non-keyable x rows, because a row that cannot match anything is exactly what an anti join asks for.

Conditions

Failures raise typed conditions -- rurl_url_join_input_error, rurl_url_join_policy_error, rurl_url_join_suffix_error, rurl_url_join_key_name_error, rurl_url_join_relationship_error, rurl_url_join_invalid_error and rurl_url_join_warning_error, all inheriting from rurl_url_join_error -- so they can be caught precisely. Messages report row positions and truncated keys, never URL content, so a credential in the input cannot leak into an error message.

See Also

get_url_key() and url_key_policy() for the identity model, and canonical_join() for the legacy join that matches on cleaned strings.

Examples

Run this code
pages <- data.frame(
  URL = c("http://example.com:80/a", "https://example.com/b",
          "http://example.com/c?", "not a url"),
  clicks = c(10, 20, 30, 40),
  stringsAsFactors = FALSE
)
meta <- data.frame(
  URL = c("http://example.com/a", "http://example.com/b",
          "http://example.com/c"),
  title = c("A", "B", "C"),
  stringsAsFactors = FALSE
)

# Only row 1 matches: `:80` is redundant under http, but http is not https,
# and a present-but-empty query is not the same resource as no query.
url_inner_join(pages, meta, by = "URL")

# Every left row survives, unmatched ones with a typed missing payload.
url_left_join(pages, meta, by = "URL")

# Rows that could not be parsed at all.
url_anti_join(pages, meta, by = "URL")

# Expose the key you matched on.
url_inner_join(pages, meta, by = "URL", key_name = "key")

# Relaxing scheme equality brings row 2 in.
url_inner_join(pages, meta, by = "URL",
               policy = url_key_policy(scheme_equality = "http_https"))

Run the code above in your browser using DataLab