Learn R Programming

tinyplot (version 0.8.0)

type_violin: Violin and sina plot types

Description

Type functions for violin plots, which are an alternative to box plots for visualizing continuous distributions (by group) in the form of mirrored densities. type_violin() draws the smooth outline of each density, while type_sina() scatters the underlying observations as points within the same outline (similar to a beeswarm plot).

Usage

type_violin(
  bw = "nrd0",
  joint.bw = c("mean", "full", "none"),
  adjust = 1,
  kernel = c("gaussian", "epanechnikov", "rectangular", "triangular", "biweight",
    "cosine", "optcosine"),
  n = 512,
  trim = FALSE,
  width = 0.9,
  lighten = TRUE,
  singletons = c("warn", "drop", "none")
)

type_sina( bw = "nrd0", joint.bw = c("mean", "full", "none"), adjust = 1, kernel = c("gaussian", "epanechnikov", "rectangular", "triangular", "biweight", "cosine", "optcosine"), n = 512, trim = FALSE, width = 0.9, method = c("quasirandom", "random"), singletons = c("keep", "warn", "drop") )

Arguments

bw

the smoothing bandwidth to be used, see density for details and options.

joint.bw

character string indicating whether (and how) the smoothing bandwidth should be computed from the joint data distribution when there are multiple subgroups. The options are "mean" (the default), "full", and "none". Also accepts a logical argument, where TRUE maps to "mean" and FALSE maps to "none". See the "Bandwidth selection" section below for a discussion of practical considerations.

adjust

the bandwidth used is actually adjust*bw. This makes it easy to specify values like ‘half the default’ bandwidth.

kernel

a character string giving the smoothing kernel to be used. This must partially match one of "gaussian", "rectangular", "triangular", "epanechnikov", "biweight", "cosine" or "optcosine", with default "gaussian", and may be abbreviated to a unique prefix (single letter).

"cosine" is smoother than "optcosine", which is the usual 'cosine' kernel in the literature and almost MSE-efficient. However, "cosine" is the version used by S.

n

the number of equally spaced points at which the density is to be estimated. When n > 512, it is rounded up to a power of 2 during the calculations (as fft is used) and the final result is interpolated by approx. So it almost always makes sense to specify n as a power of two.

trim

logical indicating whether the densities should be trimmed to the range of the data. Default is FALSE. For type_sina() this only affects the envelope that bounds the displacement, so it matters mainly for exact alignment with a type_violin() layer drawn the same way.

width

numeric (ideally in the range [0, 1], although this isn't enforced) giving the normalized width of the individual violins. For type_sina() this is the width of the (undrawn) violin that the points are scattered inside.

lighten

logical. Should the fills use a lighter, opaque tint of the series colour(s)? Default is TRUE, which keeps single- and multi-group displays consistent and lets the fill read cleanly over grid lines. Set to FALSE to use the fully-saturated palette colour(s) instead. Only applies to type_violin(), since type_sina() has no fill of its own.

singletons

character string indicating what to do with singleton groups, i.e. combinations of x, by, and facet that consist of only 1 row. Both types accept "warn", which removes any singleton cases and emits a warning reporting how many there were, and "drop", which does the same thing quietly. In either case the dropped groups may still be represented as empty violins or facets in your plot.

The remaining option differs by type, as does the default. type_violin() defaults to "warn" and also accepts "none", which skips all singleton checks and retains the affected groups; possibly leading to an error. Note that singletons then require a numeric bw, since the data-driven bandwidth rules need at least 2 observations.

type_sina() instead defaults to "keep", which draws the lone observation on its group's tick. No density can be estimated from a single point, but the point itself is still worth showing, and discarding an observation from what is fundamentally a scatter plot is worse than discarding an unrenderable violin. There is no "none" for this type, since "keep" already retains these cases without error.

method

character string giving how type_sina() spreads points across the available width at each y value. "quasirandom" (the default) walks a low-discrepancy sequence, which fills the width more evenly than chance does and---unlike type_jitter---is deterministic, so repeated calls give the same plot without setting a seed. "random" draws the displacements uniformly at random instead.

Details

See type_density for more details and considerations related to bandwidth selection and kernel types.

A sina plot (Sidiropoulos et al., 2018) is closely related to a beeswarm plot, but is arguably the more principled of the two. Both spread a group's observations sideways to expose its shape. A beeswarm does so by packing points until they no longer collide, which makes its width an artefact of the rendering: change the symbol size or the device and the swarm changes shape. A sina instead displaces each point by the kernel density at its own y value, so its width is a property of the data, stable across devices and comparable between groups. The trade-off is occlusion: a beeswarm guarantees that no point hides another, whereas a sina accepts the occasional overlap. If you need collision-free packing, use a dedicated package such as beeswarm.

References

Sidiropoulos, N., Sohi, S. H., Pedersen, T. L., Porse, B. T., Winther, O., Rapin, N., and Bagger, F. O. (2018). SinaPlot: An Enhanced Chart for Simple and Truthful Representation of Single Observations Over Multiple Classes. Journal of Computational and Graphical Statistics, 27(3), 673-676. Available: https://doi.org/10.1080/10618600.2017.1366914

See Also

type_boxplot, type_density and type_ridge for the other ways of displaying a distribution by group, and type_jitter for displacing points without reference to a density.

Examples

Run this code
# "violin" type convenience string
tinyplot(weight ~ feed, data = chickwts, type = "violin")

# to match the defaults of `ggplot2::geom_violin()`, use `trim = TRUE` and
# `joint.bw = FALSE`
tinyplot(
  weight ~ feed, data = chickwts,
  # type = type_violin(trim = TRUE, joint.bw = FALSE) # same but see final ex.
  type = "violin", trim = TRUE, joint.bw = FALSE
)

# For flipped violin plots, it's usually better to use a dynamic theme to
# accommodate (horizontal) y-axis labels
tinyplot(
  weight ~ feed, data = chickwts, type = "violin", flip = TRUE,
  theme = "dynamic" # or "clean(2)", "classic", "minimal", etc.
)

# you can group by the x var to add colour (here with the original orientation)
tinyplot(weight ~ feed | feed, data = chickwts, type = "violin", legend = FALSE)

# dodged grouped violin plot example (different dataset)
tinyplot(len ~ dose | supp, data = ToothGrowth, type = "violin")

# the "sina" type shows the observations themselves, rather than a smooth
# outline drawn around them
tinyplot(weight ~ feed, data = chickwts, type = "sina")

# layering a sina on top of a violin lines up exactly, since the points are
# displaced by the violin's own half-width
tinyplot(weight ~ feed, data = chickwts, type = "violin")
tinyplot_add(type = "sina", pch = 16, col = "black")

# unlike `type_violin()`, `type_sina()` supports a continuous `by` variable;
# it colours the points rather than splitting them into groups
tinyplot(
  Sepal.Length ~ Species | Petal.Width, data = iris,
  type = "sina", pch = 16
)

# note: above we relied on `...` argument passing alongside the type
# convenience strings. But this won't work for `width`, since it will
# clash with the top-level `tinyplot(..., width = )` arg. To ensure
# correct arg passing, it's safer to use the functional type.
tinyplot(
  len ~ dose | supp, data = ToothGrowth,
  type = type_violin(width = 0.75)
)

Run the code above in your browser using DataLab