Learn R Programming

rtables (version 0.6.17)

split_cols_by_cuts: Split on static or dynamic cuts of the data

Description

Create columns (or row splits) based on values (such as quartiles) of var.

Usage

split_cols_by_cuts(
  lyt,
  var,
  cuts,
  cutlabels = NULL,
  split_label = var,
  nested = TRUE,
  cumulative = FALSE,
  show_colcounts = FALSE,
  colcount_format = NULL
)

split_rows_by_cuts( lyt, var, cuts, cutlabels = NULL, split_label = var, parent_name = var, format = NULL, na_str = NA_character_, nested = TRUE, at_sibling = NULL, cumulative = FALSE, label_pos = if (!is.null(at_sibling)) "visible" else "default", section_div = NA_character_ )

split_cols_by_cutfun( lyt, var, cutfun = qtile_cuts, cutlabelfun = function(x) NULL, split_label = var, nested = TRUE, extra_args = list(), cumulative = FALSE, show_colcounts = FALSE, colcount_format = NULL )

split_cols_by_quartiles( lyt, var, split_label = var, nested = TRUE, extra_args = list(), cumulative = FALSE, show_colcounts = FALSE, colcount_format = NULL )

split_rows_by_quartiles( lyt, var, split_label = var, parent_name = var, format = NULL, na_str = NA_character_, nested = TRUE, at_sibling = NULL, child_labels = c("default", "visible", "hidden"), extra_args = list(), cumulative = FALSE, indent_mod = 0L, label_pos = if (!is.null(at_sibling)) "visible" else "default", section_div = NA_character_ )

split_rows_by_cutfun( lyt, var, cutfun = qtile_cuts, cutlabelfun = function(x) NULL, split_label = var, parent_name = var, format = NULL, na_str = NA_character_, nested = TRUE, at_sibling = NULL, child_labels = c("default", "visible", "hidden"), extra_args = list(), cumulative = FALSE, indent_mod = 0L, label_pos = if (!is.null(at_sibling)) "visible" else "default", section_div = NA_character_ )

Value

A PreDataTableLayouts object suitable for passing to further layouting functions, and to build_table().

Arguments

lyt

(PreDataTableLayouts)
layout object pre-data used for tabulation.

var

(string)
variable name.

cuts

(numeric)
cuts to use.

cutlabels

(character or NULL)
labels for the cuts.

split_label

(string)
label to be associated with the table generated by the split. Not to be confused with labels assigned to each child (which are based on the data and type of split during tabulation).

nested

(logical)
whether this layout instruction should be applied within the existing layout structure if possible (TRUE, the default) or as a new top-level element (FALSE). Ignored if it would nest a split underneath analyses, which is not allowed.

cumulative

(flag)
whether the cuts should be treated as cumulative. Defaults to FALSE.

show_colcounts

(logical(1))
should column counts be displayed at the level facets created by this split. Defaults to FALSE.

colcount_format

(character(1))
if show_colcounts is TRUE, the format which should be used to display column counts for facets generated by this split. Defaults to "(N=xx)".

parent_name

(character(1))
Name to assign to the table corresponding to the split or group of sibling analyses, for split_rows_by* and analyze* when analyzing more than one variable, respectively. Ignored when analyzing a single variable.

format

(string, function, or list)
format associated with this split. Formats can be declared via strings ("xx.x") or function. In cases such as analyze calls, they can be character vectors or lists of functions. See formatters::list_valid_format_labels() for a list of all available format strings.

na_str

(string)
string that should be displayed when the value of x is missing. Defaults to "NA".

at_sibling

(character(1) or NULL)
If non-null, a preceding split or analyze to anchor this instruction to as a direct sibling. Cannot select an instruction that is downstream of a point where a previously used anchor (See Nesting Anchor Resolution for details).

label_pos

(string)
location where the variable label should be displayed. Accepts "hidden" (default for non-analyze row splits), "visible", "topleft", and "default" (for analyze splits only). For analyze calls, "default" indicates that the variable should be visible if and only if multiple variables are analyzed at the same level of nesting.

section_div

(string)
string which should be repeated as a section divider after each group defined by this split instruction, or NA_character_ (the default) for no section divider.

cutfun

(function)
function which accepts the full vector of var values and returns cut points to be used (via cut) when splitting data during tabulation.

cutlabelfun

(function)
function which returns either labels for the cuts or NULL when passed the return value of cutfun.

extra_args

(list)
extra arguments to be passed to the tabulation function. Element position in the list corresponds to the children of this split. Named elements in the child-specific lists are ignored if they do not match a formal argument of the tabulation function.

child_labels

(string)
the display behavior for the labels (i.e. label rows) of the children of this split. Accepts "default", "visible", and "hidden". Defaults to "default" which flags the label row as visible only if the child has 0 content rows.

indent_mod

(numeric)
modifier for the default indent position for the structure created by this function (subtable, content table, or row) and all of that structure's children. Defaults to 0, which corresponds to the unmodified default behavior.

Nesting Anchor Resolution

When nested is TRUE, at_sibling allows you to set a nesting anchor that your new split_rows_by* or analyze* directive should be placed as a sibling to. The lookup for this anchor occurs only in the currently active top-level nesting stack, meaning the directives that have occurred since the last split or analysis with nested == FALSE.

Furthermore, resolution occurs against the first element of each arm of a branching point caused by any previous uses of at_sibling but only descends into the last arm.

So for example if our previous layout was generated via:

lyt <- basic_table() |>
  split_rows_by("SEX") |>
  analyze("AGE") |>
  split_rows_by("BMRKR2", nested = FALSE) |>
  split_rows_by("RACE") |>
  analyze("AGE") |>
  split_rows_by("SEX", at_sibling = "RACE") |>
  analyze("BMRKR1")

The eligible anchor points would be "BMRKR2", "RACE", "SEX" and "BMRKR1". "AGE" is masked by the branching caused by anchoring our SEX split on RACE.

Finally, while at_sibling does support de-duplication of "<name>[i]" anchors, it does so within the set of available anchors, which can be counter-intuitive. It is strongly suggested that the parent_name and table_names argument(s) of split_rows_by* and analyze be used to prevent the need for this. at_sibling will resolve to table names overridden in this manner.

Author

Gabriel Becker

Details

For dynamic cuts, the cut is transformed into a static cut by build_table() based on the full dataset, before proceeding. Thus even when nested within another split in column/row space, the resulting split will reflect the overall values (e.g., quartiles) in the dataset, NOT the values for subset it is nested under.

Examples

Run this code
library(dplyr)

# split_cols_by_cuts
lyt <- basic_table() |>
  split_cols_by("ARM") |>
  split_cols_by_cuts("AGE",
    split_label = "Age",
    cuts = c(0, 25, 35, 1000),
    cutlabels = c("young", "medium", "old")
  ) |>
  analyze(c("BMRKR2", "STRATA2")) |>
  append_topleft("counts")

tbl <- build_table(lyt, ex_adsl)
tbl

# split_rows_by_cuts
lyt2 <- basic_table() |>
  split_cols_by("ARM") |>
  split_rows_by_cuts("AGE",
    split_label = "Age",
    cuts = c(0, 25, 35, 1000),
    cutlabels = c("young", "medium", "old")
  ) |>
  analyze(c("BMRKR2", "STRATA2")) |>
  append_topleft("counts")


tbl2 <- build_table(lyt2, ex_adsl)
tbl2

# split_cols_by_quartiles

lyt3 <- basic_table() |>
  split_cols_by("ARM") |>
  split_cols_by_quartiles("AGE", split_label = "Age") |>
  analyze(c("BMRKR2", "STRATA2")) |>
  append_topleft("counts")

tbl3 <- build_table(lyt3, ex_adsl)
tbl3

# split_rows_by_quartiles
lyt4 <- basic_table(show_colcounts = TRUE) |>
  split_cols_by("ARM") |>
  split_rows_by_quartiles("AGE", split_label = "Age") |>
  analyze("BMRKR2") |>
  append_topleft(c("Age Quartiles", " Counts BMRKR2"))

tbl4 <- build_table(lyt4, ex_adsl)
tbl4

# split_cols_by_cutfun
cutfun <- function(x) {
  cutpoints <- c(
    min(x),
    mean(x),
    max(x)
  )

  names(cutpoints) <- c("", "Younger", "Older")
  cutpoints
}

lyt5 <- basic_table() |>
  split_cols_by_cutfun("AGE", cutfun = cutfun) |>
  analyze("SEX")

tbl5 <- build_table(lyt5, ex_adsl)
tbl5

# split_rows_by_cutfun
lyt6 <- basic_table() |>
  split_cols_by("SEX") |>
  split_rows_by_cutfun("AGE", cutfun = cutfun) |>
  analyze("BMRKR2")

tbl6 <- build_table(lyt6, ex_adsl)
tbl6

Run the code above in your browser using DataLab