Learn R Programming

soundgen (version 3.0.0)

compareSounds: Compare sounds

Description

Computes distances between sounds based on comparing their spectrogram-like representations. compareSounds takes two sounds or feature matrices as input, whereas compareFolder takes a path to a folder with audio files or a list of feature matrices and returns a matrix of pairwise distances between them. Feature matrices are normalized and compared with Dynamic Time Warp (DTW), correlation, cosine distance, or pixel by pixel.

Usage

compareSounds(
  x,
  y,
  samplingRate = NULL,
  specFun = "melspec",
  specFun_pars = list(),
  logSpec = FALSE,
  method = c("cor", "cosine", "diff", "dtw"),
  padWith = NA,
  padDir = c("central", "left", "right"),
  dtw_pars = list()
)

compareFolder( myfolder = NULL, spectrograms = NULL, matchAllLengths = FALSE, specFun = "melspec", specFun_pars = list(), logSpec = FALSE, method = c("cor", "cosine", "diff", "dtw"), padWith = NA, padDir = c("central", "left", "right"), dtw_pars = list(), cores = 1, reportEvery = NULL )

Value

compareSounds returns a dataframe with two columns: "method" for the method(s) used, and "distance" for the distance between the two sounds calculated with that method. The range of distances is [0, 1]

for "cor", "cosine", and "diff", and [0, Inf) for "dtw".

compareFolder returns a list of distance matrices (dist

objects), one for each method.

Arguments

x, y

either two matrices (spectrograms or feature matrices) or two sounds to be compared (numeric vectors, Wave objects, or paths to wav/mp3 files)

samplingRate

if one or both inputs are numeric vectors, specify sampling rate, Hz. This does not resample audio. For meaningful comparisons, audio inputs should already have the same sampling rate

specFun

the function used to extract a spectrogram-like feature matrix. Can be a string or a custom function that takes audio (numeric vector) as the first argument and returns a spectrogram-like matrix with time in columns and features in rows (see examples). Supported strings:

stft_simple

'stft' / STFT / 'stft_simple' (amplitude spectrogram) Parameters in specFun_pars: samplingRate (optional if frequency and time labels are not needed), wl (samples), step (samples), wn, zp, padWithSilence.

spectrogram

'spectrogram' (amplitude spectrogram - a wrapper around stft_simple with more options). Parameters in specFun_pars: see spectrogram.

powspec

'powerspec' (power spectrogram with tuneR). Parameters in specFun_pars: wintime (s), steptime (s), dither.

melfcc

'melspec' (mel-spectrogram with tuneR), 'mfcc' / 'melfcc' (MFCCs). Parameters in specFun_pars: windowLength (ms), step (ms), nbands, maxfreq (Hz), MFCC (integer vector).

audSpectrogram

'audSpectrogram' / 'audSpec' (auditory spectrogram). Parameters in specFun_pars: see audSpectrogram.

getRMS

'getRMS' / 'rms' (RMS amplitude envelope). Parameters in specFun_pars: see getRMS.

getEnv

'getEnv' / 'env' (upsampled envelopes: RMS, analytical, peak, etc.). Parameters in specFun_pars: see getEnv.

spectrum

'spectrum' (short-term spectrum). Parameters in specFun_pars: see spectrum.

meanSpectrum

'meanSpectrum' / 'meanspec' / 'meanSpec' (long-term average spectrum). Parameters in specFun_pars: see meanSpectrum.

ssm

'ssm' (self-similarity matrix). Parameters in specFun_pars: see ssm.

modulationSpectrum

'ms' / 'modulationSpectrum' (modulation spectrum). Parameters in specFun_pars: see modulationSpectrum.

specFun_pars

a list of parameters passed to specFun

logSpec

if TRUE, applies a log transform to the spectrograms before normalization

method

method(s) of comparing spectrograms of two sounds: "cor" = Pearson's correlation distance; "cosine" = cosine distance; "diff" = normalized absolute difference; "dtw" = multivariate Dynamic Time Warp with dtw (NB: the "dtw" package must be installed for this method to work)

padWith

if the durations of x and y are not identical, the compared spectrograms are either padded with silence (padWith = 0) or truncated (padWith = NA) to have the same number of columns. Padding with NA is like truncating the longer sound to the short one's duration, whereas padding with 0 means that the shorter sound is padded with zero to the long one's duration

padDir

if padding, specify where to add zeros or NAs: before the sound ('left'), after the sound ('right'), or on both sides ('central')

dtw_pars

a list of parameters passed to dtw

myfolder

path to folder containing audio files to compare

spectrograms

a list of spectrogram-like feature matrices to use instead of analyzing the audio - extracted acoustic features, modulation spectra, similarity matrices, ... (overrides myfolder)

matchAllLengths

if TRUE, all spectrograms are length-matched - e.g., if 100 sounds are compared and padWith = 0, 99 of them are padded with silence to match the longest sound (faster); if FALSE, length matching is performed separately for each compared pair. Note that this also affects DTW matching if other distance metrics are used at the same time; thus, you may want to disable matchAllLengths if using DTW

cores

number of cores for parallel processing

reportEvery

when processing multiple inputs, report estimated time left every reportEvery iterations (NULL = default, NA = don't report); see reportTime

Details

If the input is audio, several methods of producing spectrograms are available ("specFun"). For more customized options, just prepare your spectrograms or feature matrices first (time in columns, features like pitch, peak frequency, etc. in rows), and then pass them to compareSounds (see examples). All methods except for DTW require that the compared matrices should be of the same size. Compared sounds should ideally have the same sampling rate. If they differ, row (frequency bin) truncation is performed by position, keeping only the first min(nrow) rows, which approximates keeping frequencies up to the lower Nyquist frequency when both spectrograms use the same frequency resolution. In case of differences in duration, the shorter sound is padded with 0 (silence) or NA, as controlled by arguments padWith, padDir. If passing custom feature matrices, ensure that they have the same dimensions or that padding with 0 (silence) makes sense.

Examples

Run this code
s1 = soundgen(sylLen = 100, pitch = c(80, 180), formants = 'a')
s2 = soundgen(sylLen = 120, pitch = c(150, 350), formants = 'u')
compareSounds(s1, s2, samplingRate = 16000, method = c('cor', 'cosine', 'diff'))
# spectrogram(s1); playme(s1)
# spectrogram(s2); playme(s2)

if (FALSE) {
# NB: install the "dtw" library to run the examples

# compare all sounds in a folder, e.g.:
target = '~/Documents/Research/zz_test_audio/temp_long'
cf = compareFolder(target, logSpec = TRUE)
mds = as.data.frame(cmdscale(cf$cor))
plot(mds, type = 'n'); text(mds, labels = abbreviate(rownames(mds)))

# or use manually produced spectrograms
sp = spectrogram(target, windowLength = c(10, 40), overlap = 75,
  yScale = 'ERB', output = 'processed', plot = FALSE, cores = 4)
image(sp[[1]])
cf1 = compareFolder(spectrograms = sp)
mds1 = as.data.frame(cmdscale(cf1$cor))
plot(mds1, type = 'n'); text(mds1, labels = abbreviate(rownames(mds1)))

# extract a spectrogram-like representation using a custom function
# (e.g., full-resolution analytic envelopes instead of downsampled RMS)
compareSounds(s1, s2, samplingRate = 16000,
  specFun = function(x) matrix(hilbert_approx(x)$envelope, nrow = 1))

# some more examples
s1 = soundgen(formants = 'a', play = TRUE)
s2 = soundgen(formants = 'ae', play = TRUE)
s3 = soundgen(formants = 'eae', sylLen = 700, play = TRUE)
s4 = runif(8000, -1, 1)  # white noise
compareSounds(s1, s2, samplingRate = 16000)
compareSounds(s1, s4, samplingRate = 16000)

# the central section of s3 is more similar to s1 than is the beg/end of s3
compareSounds(s1, s3, samplingRate = 16000, padDir = 'left')
compareSounds(s1, s3, samplingRate = 16000, padDir = 'central')

# padding with 0 penalizes differences in duration, whereas padding with NA
# is like saying we only care about the overlapping part
compareSounds(s1, s4, samplingRate = 16000, padWith = 0)
compareSounds(s1, s4, samplingRate = 16000, padWith = NA)

# different types of spectrograms produce quite different results
compareSounds(s1, s3, samplingRate = 16000, specFun = 'stft')
compareSounds(s1, s3, samplingRate = 16000, specFun = 'melspec')
compareSounds(s1, s3, samplingRate = 16000, specFun = 'mfcc')
compareSounds(s1, s3, samplingRate = 16000, specFun = 'audSpec')

# pass additional control parameters to specFun and DTW
compareSounds(s1, s3, samplingRate = 16000,
              specFun = 'melspec',
              specFun_pars = list(nbands = 128),
              dtw_pars = list(dist.method = "Manhattan"))

# use feature matrices instead of spectrograms
# (time in columns, features in rows)
a1 = t(as.matrix(analyze(s1, samplingRate = 16000)$detailed))
a1 = a1[4:nrow(a1), ]; a1[is.na(a1)] = 0  # don't use dur and time stamps
a2 = t(as.matrix(analyze(s2, samplingRate = 16000)$detailed))
a2 = a2[4:nrow(a2), ]; a2[is.na(a2)] = 0
a4 = t(as.matrix(analyze(s4, samplingRate = 16000)$detailed))
a4 = a4[4:nrow(a4), ]; a4[is.na(a4)] = 0
compareSounds(a1, a2, method = c('cosine', 'dtw'))
compareSounds(a1, a4, method = c('cosine', 'dtw'))
}

Run the code above in your browser using DataLab