Performs a join between two data frames by canonicalizing URLs to a shared
"clean" format using safe_parse_urls and then matching on
that key.
This is suitable for large crawl exports.
canonical_join(
data_A,
data_B,
col_A = "URL",
col_B = "URL",
suffix_A = "_A",
suffix_B = "_B",
name_A = NULL,
name_B = NULL,
join = c("inner", "left", "right", "full"),
collision = c("first", "all", "error"),
on_parse_error = c("keep", "drop", "error"),
join_parse_status = c("ok", "ok_or_warning"),
...
)A data frame representing the join. The output includes:
The original URL columns (named via name_A / name_B,
or after the input expressions when those are NULL).
JoinKey: the canonicalized URL used for matching.
All other columns from data_A and data_B with
suffixes applied.
Returns an empty data frame with the expected structure if no matches are found or if inputs are invalid.
A data frame containing URLs for the left side of the join.
A data frame containing URLs for the right side of the join.
Character string, the name of the column in data_A that
contains URLs. Defaults to "URL".
Character string, the name of the column in data_B that
contains URLs. Defaults to "URL".
Character string, suffix to append to data_A columns
(excluding the URL column) in the output. Defaults to "_A".
Character string, suffix to append to data_B columns
(excluding the URL column) in the output. Defaults to "_B".
Character string, the name of the output column holding the
original data_A URLs. Defaults to NULL, in which case the
name is derived from the data_A argument expression via
deparse(substitute()). Supply an explicit value for stable output
names when piping or passing anonymous inputs (e.g.
canonical_join(df[df$x > 1, ], get_b())).
Character string, the name of the output column holding the
original data_B URLs. Defaults to NULL; behaves like
name_A for data_B.
Join type: "inner", "left", "right", or
"full". Defaults to "inner".
How to handle duplicate canonical keys within inputs.
"first" keeps the first row per key, "all" keeps all rows
(many-to-many), and "error" stops on duplicates. Defaults to
"first".
How to handle URLs that fail canonicalization.
"keep" retains them as unmatched rows (for left/right/full joins),
"drop" removes them before joining, and "error" stops.
Defaults to "keep".
Which parse statuses yield joinable canonical keys.
"ok" (default) joins only rows whose parse_status begins with
"ok" ("ok", "ok-ftp", "ok-scheme-relative").
"ok_or_warning" additionally treats parseable-but-suspicious
warning-* statuses ("warning-no-tld",
"warning-invalid-tld", "warning-public-suffix") as joinable.
Joining on warning statuses can increase false-positive matches between
distinct hosts that both fail TLD derivation.
Additional arguments forwarded to safe_parse_urls,
controlling canonicalization (e.g., protocol_handling,
www_handling, trailing_slash_handling,
index_page_handling, path_normalization,
scheme_relative_handling, host_encoding,
path_encoding, the url_standard selector, and the
profile bundle). When url_standard is set, forwarding a
governed low-level knob it would override (e.g. path_normalization)
is an error, exactly as in safe_parse_url; the orthogonal
path_encoding and host_encoding presentation knobs layer
freely on any profile. A profile (e.g. "seo",
"whatwg") may also be forwarded: like safe_parse_url,
it bundles several knobs, expands only into knobs you did not supply, and
an explicit knob always overrides it (so the url_standard conflict check is
skipped on the profile path). Inspect a bundle with
url_profile. See "Legacy presentation dials" below for the
arguments that warn.
canonical_join() keys the join on the cleaned presentation string
(clean_url), so every cleaning or display argument forwarded through
... currently changes which rows match. Those arguments do not
participate in URL identity; they are legacy behavior retained for a
deprecation window. Supplying any of
protocol_handling, www_handling, source,
tld_source, case_handling, trailing_slash_handling,
index_page_handling, path_normalization,
subdomain_levels_to_keep, host_encoding, path_encoding,
port_handling, engine, profile, or any query cleaning
dial (query_handling, params_keep, params_drop,
params_case_sensitive, sort_params,
empty_param_handling, decode_plus)
emits one warning per call, of class
"rurl_legacy_join_dial_warning". Results are unchanged: the warning
is purely additive, so no caller is silently re-matched.
The input and interpretation arguments url_standard,
scheme_acceptance, scheme_policy, and
scheme_relative_handling are legitimate inputs to identity and never
warn.
Because the condition is classed, it can be silenced selectively without
hiding other warnings:
suppressWarnings(canonical_join(A, B, www_handling = "strip"),
classes = "rurl_legacy_join_dial_warning").
A <- data.frame(
URL = c("https://Example.com/page", "https://example.com/other"),
ValA = 1:2, stringsAsFactors = FALSE
)
B <- data.frame(
URL = c(
"https://example.com/page?utm_source=nl",
"https://example.com/missing"
),
ValB = c("x", "y"), stringsAsFactors = FALSE
)
# Default canonicalization lower-cases the host and drops the query, so the
# first row of each side shares one canonical key.
canonical_join(A, B)
Run the code above in your browser using DataLab