Learn R Programming

soundgen (version 3.0.0)

ssm: Self-similarity matrix

Description

Calculates the self-similarity matrix and novelty vector of a sound. The self-similarity matrix is produced by comparing all pairs of frames in the input sound. Novelty is calculated by convolving the self-similarity matrix with a tapered checkerboard kernel. The positive lobes of the kernel represent coherence (self-similarity within the regions on either side of the center point) and the negative lobes anti-coherence (cross-similarity between these two regions). Since novelty is the dot product of the checkerboard kernel with the SSM, it is high when the two regions are self-similar (internally consistent) but different from each other.

Usage

ssm(
  x,
  samplingRate = NULL,
  from = NULL,
  to = NULL,
  specFun = "melspec",
  specFun_pars = list(),
  logSpec = FALSE,
  normalize = FALSE,
  simil = c("cosine", "cor"),
  kernelLen = 1000,
  kernelSD = 0.5,
  padWith = 0,
  ssmWin = 1,
  summaryFun = c("mean", "sd"),
  output = c("ssm", "novelty"),
  reportEvery = NULL,
  cores = 1,
  plot = TRUE,
  savePlots = FALSE,
  embed = FALSE,
  main = NULL,
  heights = c(2, 1),
  width = 900,
  height = 500,
  units = "px",
  res = NA,
  specPars = list(colorTheme = c("bw", "seewave", "heat.colors", "...")[2], xlab =
    "Time"),
  ssmPars = list(colorTheme = c("bw", "seewave", "heat.colors", "...")[2], xlab = "Time",
    ylab = "Time"),
  noveltyPars = list(type = "l", col = "black", lwd = 1)
)

Value

A list with two top-level elements: $detailed and

$summary.

$detailed contains per-file results selected with the output

argument. If multiple sounds are analyzed, $detailed is a named list of per-sound lists. If a single sound is analyzed, it is simplified to a single list. Each list may contain:

ssm

self-similarity matrix

novelty

novelty vector

$summary is a dataframe with summaries of novelty per sound (one row per file), or NULL if summaryFun = NULL.

Arguments

x

path to a folder, one or more wav or mp3 files c('file1.wav', 'file2.mp3'), Wave object, numeric vector, or a list of Wave objects or numeric vectors

samplingRate

sampling rate of x (only needed if x is a numeric vector)

from, to

if NULL (default), analyzes the whole sound, otherwise from...to (s)

specFun

the function used to extract a spectrogram-like feature matrix. Can be a string or a custom function that takes audio (numeric vector) as the first argument and returns a spectrogram-like matrix with time in columns and features in rows (see examples). A precomputed matrix is also accepted (features in rows, time [ms] in columns, numeric rownames for plotting). Supported strings:

stft_simple

'stft' / STFT / 'stft_simple' (amplitude spectrogram) Parameters in specFun_pars: samplingRate (optional if frequency and time labels are not needed), wl (samples), step (samples), wn, zp, padWithSilence.

spectrogram

'spectrogram' (amplitude spectrogram - a wrapper around stft_simple with more options). Parameters in specFun_pars: see spectrogram.

powspec

'powerspec' (power spectrogram with tuneR). Parameters in specFun_pars: wintime (s) or windowLength (ms), steptime (s) or step (ms), dither.

melfcc

'melspec' (mel-spectrogram), 'mfcc' / 'melfcc' (MFCCs). Parameters in specFun_pars: windowLength (ms), step (ms), nbands, maxfreq (Hz), MFCC (integer vector).

audSpectrogram

'audSpectrogram' / 'audSpec' (auditory spectrogram). Parameters in specFun_pars: see audSpectrogram.

getRMS

'getRMS' / 'rms' (RMS amplitude envelope). Parameters in specFun_pars: see getRMS.

getEnv

'getEnv' / 'env' ("various smoothed envelopes: RMS, analytical, peak, etc.). Parameters in specFun_pars: see getEnv.

specFun_pars

a list of parameters passed to specFun. Defaults for specFun = "melspec": list(windowLength = 125, step = 25, nbands = 50); defaults for specFun = "audSpec": list(nFilters = 16, step = 10)

logSpec

if TRUE, the input is log-transformed prior to calculating self-similarity

normalize

if TRUE, the spectrum of each STFT frame (or each column of feature matrix) is normalized to the same range prior to calculating self-similarity

simil

method for comparing frames: "cosine" = cosine similarity, "cor" = Pearson's correlation

kernelLen

length of checkerboard kernel for calculating novelty, ms (larger values favor global, slow vs. local, fast novelty)

kernelSD

SD of checkerboard kernel evaluated over [-1, 1]: for ex., if kernelSD = 0.5, the kernel spans approximately ±2 SDs

padWith

how to treat edges when calculating novelty: NA = pad with NA (ignores edges in correlation), 0 = pad with zeros

ssmWin

window for averaging SSM, frames (has a smoothing effect and speeds up the processing)

summaryFun

functions used to summarize novelty, eg c('mean', 'sd'); user-defined functions are fine (see examples); NAs are omitted automatically for mean/median/sd/min/max/range/sum, otherwise take care of NAs yourself

output

what to include in $detailed (drop "ssm" to save memory when analyzing a lot of files); options: 'ssm', 'novelty', 'all'

reportEvery

when processing multiple inputs, report estimated time left every reportEvery iterations (NULL = default, NA = don't report); see reportTime

cores

number of cores for parallel processing

plot

if TRUE, plots the SSM

savePlots

if TRUE, creates a subdirectory in the input directory (if input is a file or folder) or in the working directory (if input is a vector etc), named after the function (eg "spectrogram/"). All plots and audio files (if any) are saved in this new directory. If there are multiple inputs, an html notebook is also created for easy viewing and listening

embed

if TRUE and savePlots is set and there are multiple inputs, all saved images and audio (if any) are embedded in the exported html notebook for easy sharing; if FALSE, the html file links to separate images and audio files (but separate files are still saved). NB: for this to work, package "base64enc" must be installed

main

plot title

heights

relative sizes of the SSM and spectrogram/novelty plot

width, height, units, res

graphical parameters for saving plots passed to png

specPars

graphical parameters passed to filled.contour.mod and affecting the spectrogram

ssmPars

graphical parameters passed to filled.contour.mod and affecting the plot of SSM

noveltyPars

graphical parameters passed to lines and affecting the novelty contour

References

  • Foote, J. (1999, October). Visualizing music and audio using self-similarity. In Proceedings of the seventh ACM international conference on Multimedia (Part 1) (pp. 77-80). ACM.

  • Foote, J. (2000). Automatic audio segmentation using a measure of audio novelty. In Multimedia and Expo, 2000. ICME 2000. 2000 IEEE International Conference on (Vol. 1, pp. 452-455). IEEE.

See Also

spectrogram modulationSpectrum segment

Examples

Run this code
sound = c(soundgen(),
          soundgen(nSyl = 4, sylLen = 50, pauseLen = 70,
          formants = NA, pitch = c(500, 330)))
# playme(sound)
# detailed, local features (captures each syllable)
s1 = ssm(sound, samplingRate = 16000, kernelLen = 100)
# more global features (captures the transition b/w the two sounds)
s2 = ssm(sound, samplingRate = 16000, kernelLen = 400)

s2$summary
s2$detailed$novelty  # novelty contour
if (FALSE) {
ssm(sound, samplingRate = 16000,
    specFun = 'mfcc', simil = 'cor', normalize = TRUE,
    ssmWin = 10,  # speed up the processing
    kernelLen = 300,  # global features
    specPars = list(colorTheme = 'seewave'),
    ssmPars = list(col = rainbow(100)),
    noveltyPars = list(type = 'l', lty = 3, lwd = 2))

# Custom input: produce a nice spectrogram first, then feed it into ssm()
sp = spectrogram(sound, 16000, windowLength = c(5, 40), contrast = .3,
  output = 'processed')  # return the modified spectrogram
ssm(sound, 16000, kernelLen = 400, specFun = sp)

# Custom input: use acoustic features returned by analyze()
an = analyze(sound, 16000, windowLength = 20, novelty = NULL)
feature_mat = t(an$detailed[, 4:ncol(an$detailed)]) # or select pitch, HNR, ...
feature_mat = t(apply(feature_mat, 1, scale))  # z-transform all variables
feature_mat[is.na(feature_mat)] = 0  # get rid of NAs
colnames(feature_mat) = an$detailed$time  # time stamps in ms
rownames(feature_mat) = 1:nrow(feature_mat)
image(t(feature_mat))  # not a spectrogram, just a feature matrix
ssm(sound, 16000, kernelLen = 500, specFun = feature_mat, logSpec = FALSE,
  specPars = list(ylab = 'Feature'))
}

Run the code above in your browser using DataLab