Learn R Programming

soundgen (version 3.0.0)

audSpectrogram: Auditory spectrogram

Description

Produces an auditory spectrogram by convolving the sound with a bank of bandpass filters. The main difference from STFT is that we don't window the signal and de facto get variable temporal resolution in different frequency channels, as with a wavelet transform. The key settings are filterType, nFilters_oct, and yScale, which determine the type, number, and spacing of the filters, respectively. Gammatone filters were designed as a simple approximation of human perception - see Slaney 1993 "An Efficient Implementation of the Patterson–Holdsworth Auditory Filter Bank". Butterworth or Chebyshev filters are not meant to model perception, but can be useful for quickly plotting a sound.

Usage

audSpectrogram(
  x,
  samplingRate = NULL,
  scale = NULL,
  from = NULL,
  to = NULL,
  step = 10,
  dynamicRange = 80,
  filterType = c("gammatone", "butterworth", "chebyshev"),
  envelope = c("rms", "hil"),
  nFilters_oct = 6,
  nFilters = NULL,
  yScale = c("ERB", "bark", "mel", "log"),
  filterOrder = NULL,
  bandwidth = NULL,
  bandwidthMult = 1,
  minFreq = 20,
  maxFreq = NULL,
  minBandwidth = 10,
  output = c("all", "audSpec", "audSpec_processed", "filterbank", "filterbank_env",
    "filters"),
  reportEvery = NULL,
  cores = 1,
  plot = TRUE,
  savePlots = FALSE,
  embed = FALSE,
  plotFilters = FALSE,
  osc = c("linear", "dB", "none"),
  heights = c(3, 1),
  ylim = NULL,
  contrast = 0,
  brightness = 0,
  maxPoints = c(1e+05, 5e+05),
  colorTheme = "bw",
  col = NULL,
  extraContour = NULL,
  xlab = NULL,
  ylab = NULL,
  xaxp = NULL,
  mar = c(5.1, 4.1, 4.1, 2),
  main = NULL,
  grid = NULL,
  width = 900,
  height = 500,
  units = "px",
  res = NA,
  ...
)

Value

A list for each analyzed file, including:

audSpec

auditory spectrogram: a matrix with frequency in rows (kHz) and time in columns (ms), offset by step/2

audSpec_processed

same dimensions, rescaled for plotting (log-transformed, contrast/brightness-adjusted, range 0–1)

filterbank

raw filter outputs: a matrix with one row per filter (ordered by center frequency) and one column per audio sample

filterbank_env

Hilbert envelopes of the filterbank, same dimensions as filterbank; NA if envelope = "rms"

filters

a dataframe giving the center frequencies, bandwidths, and lower/upper bounds of the used filters, all in Hz

Arguments

x

path to a folder, one or more wav or mp3 files c('file1.wav', 'file2.mp3'), Wave object, numeric vector, or a list of Wave objects or numeric vectors

samplingRate

sampling rate of x (only needed if x is a numeric vector)

scale

maximum possible amplitude of input, used to normalize the input vector (only needed if x is a numeric vector)

from, to

if NULL (default), analyzes the whole sound, otherwise from...to (s)

step

step, ms (determines time resolution of the plot, but not of the returned envelopes per channel). step = NULL means no downsampling at all when envelope = "hil" (ncol of output = length of input audio) and a default of 5 ms when envelope = "rms"

dynamicRange

regions under -dynamicRange dB are treated as silent

filterType

"butterworth" = Butterworth filter (IIR) butter, "chebyshev" = Chebyshev filter (IIR) cheby1, "gammatone" = gammatone filter (FIR)

envelope

the method of computing the envelope of each channel: "rms" = root mean square per window, which is faster but gives limited time resolution (default), "hil" = analytic envelope obtained with a Hilbert transform, low-pass filtered and downsampled unless step = NULL, which is slower but gives the best possible time resolution. As a simple heuristic, you may want to use "hil" if your desired time step is smaller than ~5 ms

nFilters_oct

the approximate number of filters per octave between minFreq and maxFreq; the actual resolution depends on yScale: for instance, if yScale = 'ERB', center frequencies are equally spaced on the ERB scale (set plotFilters = TRUE to check)

nFilters

an alternative way to specify frequency resolution: if specified, overrides nFilters_oct

yScale

determines the location of center frequencies of the filters

filterOrder

filter order (defaults to 4 for gammatones, 3 otherwise)

bandwidth

filter bandwidth, octaves; if NULL, defaults to ERB bandwidths

bandwidthMult

a scaling factor for all bandwidths (1 = no effect)

minFreq, maxFreq

the range of frequencies to analyze. If the spectrogram looks empty, try increasing minFreq - the lowest filters are prone to returning very large values, which can make the rest of the spectrogram look empty

minBandwidth

minimum filter bandwidth, Hz (otherwise filters may become too narrow when nFilters is high); only affects Butterworth and Chebyshev filters, not gammatones

output

character vector specifying which measures to return. Defaults to everything, but this takes a lot of RAM, so shorten to what's needed if analyzing many files at once

reportEvery

when processing multiple inputs, report estimated time left every reportEvery iterations (NULL = default, NA = don't report); see reportTime

cores

number of cores for parallel processing

plot

if TRUE, produces a plot of the results

savePlots

if TRUE, creates a subdirectory in the input directory (if input is a file or folder) or in the working directory (if input is a vector etc), named after the function (eg "spectrogram/"). All plots and audio files (if any) are saved in this new directory. If there are multiple inputs, an html notebook is also created for easy viewing and listening

embed

if TRUE and savePlots is set and there are multiple inputs, all saved images and audio (if any) are embedded in the exported html notebook for easy sharing; if FALSE, the html file links to separate images and audio files (but separate files are still saved). NB: for this to work, package "base64enc" must be installed

plotFilters

if TRUE, plots the filters as central frequencies ± bandwidth/2

osc

"none" = no oscillogram; "linear" = on the original scale; "dB" = in decibels

heights

a vector of length two specifying the relative height of the spectrogram and the oscillogram (including time axes labels)

ylim

frequency range to plot, kHz (defaults to 0 to Nyquist frequency). NB: still in kHz, even if yScale = bark, mel, or ERB

contrast

controls the sharpness or contrast of the image: <0 = decrease contrast, 0 = no change, >0 increase contrast. Recommended range approximately (-1, 1). The spectrogram is raised to the power of exp(3 * contrast)

brightness

makes the image lighter or darker, range [-1, 1] (default 0 = no change); for colorTheme = "bw", <0 = darker, >0 = lighter, range [-1, 1]. Values are remapped through a smooth sigmoid transfer curve that preserves the full color palette. To lighten or darken the palette itself, change the colors

maxPoints

the maximum number of "pixels" in the oscillogram (if any) and spectrogram; good for quickly plotting long audio files; defaults to c(1e5, 5e5); does not affect reassigned spectrograms

colorTheme

black and white ('bw'), as in seewave package ('seewave'), matlab-type palette ('matlab'), or any palette from palette such as 'heat.colors', 'cm.colors', etc

col

actual colors, eg rev(rainbow(100)) - see ?hcl.colors for colors in base R (overrides colorTheme)

extraContour

a vector of arbitrary length scaled in Hz (regardless of yScale, but nonlinear yScale also warps the contour) that will be plotted over the spectrogram (eg pitch contour); can also be a list with extra graphical parameters such as lwd, col, warp (FALSE = plot as is, TRUE = warp to conform to nonlinear yScale), etc. (see examples)

xlab, ylab, main, mar, xaxp

graphical parameters for plotting

grid

if numeric, adds n = grid dotted lines per kHz and the same number of lines along the time axis

width, height, units, res

graphical parameters for saving plots passed to png

...

other graphical parameters

Examples

Run this code
data('speechEx', package = 'soundgen')

# auditory spectrogram
asp = audSpectrogram(speechEx, to = 1, step = 5)
dim(asp$audSpec)

# compare to STFT with similar time and frequency resolution (~100 times faster)
fs = spectrogram(speechEx, to = 1, yScale = 'ERB', windowLength = 5, step = 5)
dim(fs)

if (FALSE) {
# add bells and whistles
audSpectrogram(speechEx,
  nFilters = 128,
  dynamicRange = 150,
  osc = 'none',
  heights = c(2, 1),  # spectro/osc height ratio
  contrast = .4,  # increase contrast
  brightness = -.2,  # reduce brightness
  colorTheme = 'matlab',  # pick color theme...
  # col = hcl.colors(100, palette = 'Plasma'),  # ...or specify the colors
  cex.lab = .75, cex.axis = .75,  # text size and other base graphics pars
  grid = 5,  # to customize, add manually with graphics::grid()
  ylim = c(0.05, 8),  # always in kHz
  main = 'My auditory spectrogram' # title
  # + axis labels, etc
)

# NB: frequency resolution is controlled by both nFilters and bandwidth
audSpectrogram(speechEx, to = 1, nFilters = 15, bandwidth = 1/2)
audSpectrogram(speechEx, to = 1, nFilters = 15, bandwidth = 1/10)
audSpectrogram(speechEx, to = 1, nFilters = 100, bandwidth = 1/2)
audSpectrogram(speechEx, to = 1, nFilters = 100, bandwidth = 1/10)
audSpectrogram(speechEx, to = 1, nFilters_oct = 5, bandwidth = 1/10)
audSpectrogram(speechEx, to = 1, nFilters = 200, bandwidthMult = 1/3)

# caution: if bandwidths are too narrow relative to nFilters, there may be gaps
audSpectrogram(speechEx, to = 1, nFilters = 30, bandwidthMult = 1/3,
  plotFilters = TRUE, plot = FALSE)

# different filter types
audSpectrogram(speechEx, to = 1, filterType = 'gammatone')
audSpectrogram(speechEx, to = 1, filterType = 'butterworth')
audSpectrogram(speechEx, to = 1, filterType = 'chebyshev')

# save auditory spectrograms of all audio files in a folder
audSpectrogram('~/Downloads/temp', savePlots = TRUE, cores = 4)
}

Run the code above in your browser using DataLab