Learn R Programming

soundgen (version 3.0.0)

matchPars: Match soundgen pars (experimental)

Description

Attempts to find settings for soundgen that will reproduce an existing sound. The principle is to mutate control parameters, trying to improve fit to target. The currently implemented optimization algorithm is simple hill climbing. Disclaimer: this function is experimental and may or may not work for particular tasks. It is intended as a supplement to - not replacement of - manual optimization. See https://cogsci.se/soundgen/sound_generation.html and https://cogsci.se/soundgen/matching/matching.html for more information.

Usage

matchPars(
  target,
  samplingRate = NULL,
  pars = NULL,
  specFun = "melspec",
  specFun_pars = list(),
  init = NULL,
  probMutation = 0.25,
  stepVariance = 0.1,
  maxIter = 50,
  minExpectedDelta = 0.001,
  compareSounds_pars = list(),
  verbose = TRUE,
  play = FALSE
)

Value

A list containing the history of parameters tried and their final values (pars).

Arguments

target

the sound we want to reproduce using soundgen: path to an audio file or numeric vector

samplingRate

sampling rate of target (only needed if target is a numeric vector, rather than a .wav file)

pars

arguments to soundgen that we are attempting to optimize

specFun

the function used to extract a spectrogram-like feature matrix. Can be a string or a custom function that takes audio (numeric vector) as the first argument and returns a spectrogram-like matrix with time in columns and features in rows (see examples). Supported strings:

stft_simple

'stft' / STFT / 'stft_simple' (amplitude spectrogram) Parameters in specFun_pars: samplingRate (optional if frequency and time labels are not needed), wl (samples), step (samples), wn, zp, padWithSilence.

spectrogram

'spectrogram' (amplitude spectrogram - a wrapper around stft_simple with more options). Parameters in specFun_pars: see spectrogram.

powspec

'powerspec' (power spectrogram with tuneR). Parameters in specFun_pars: wintime (s), steptime (s), dither.

melfcc

'melspec' (mel-spectrogram with tuneR), 'mfcc' / 'melfcc' (MFCCs). Parameters in specFun_pars: windowLength (ms), step (ms), nbands, maxfreq (Hz), MFCC (integer vector).

audSpectrogram

'audSpectrogram' / 'audSpec' (auditory spectrogram). Parameters in specFun_pars: see audSpectrogram.

getRMS

'getRMS' / 'rms' (RMS amplitude envelope). Parameters in specFun_pars: see getRMS.

getEnv

'getEnv' / 'env' (upsampled envelopes: RMS, analytical, peak, etc.). Parameters in specFun_pars: see getEnv.

spectrum

'spectrum' (short-term spectrum). Parameters in specFun_pars: see spectrum.

meanSpectrum

'meanSpectrum' / 'meanspec' / 'meanSpec' (long-term average spectrum). Parameters in specFun_pars: see meanSpectrum.

ssm

'ssm' (self-similarity matrix). Parameters in specFun_pars: see ssm.

modulationSpectrum

'ms' / 'modulationSpectrum' (modulation spectrum). Parameters in specFun_pars: see modulationSpectrum.

specFun_pars

a list of parameters passed to specFun

init

a list of initial values for the optimized parameters pars and the values of other arguments to soundgen that are fixed at non-default values (if any)

probMutation

the probability of a parameter mutating per iteration

stepVariance

scale factor for calculating the size of mutations

maxIter

maximum number of mutated sounds produced without improving the fit to target; iter = 0 means acoustic analysis only, no optimization

minExpectedDelta

minimum improvement in fit to target required to accept the new sound candidate

compareSounds_pars

a list of control parameters passed to compareSounds

verbose

if TRUE, reports the outcome at each iteration

play

if TRUE, plays back the accepted candidate at each iteration

Examples

Run this code
if (FALSE) {
target = soundgen(sylLen = 600, pitch = c(300, 200),
                  rolloff = -20, play = TRUE, plot = TRUE)
# we hope to reproduce this sound

# Match pars based on acoustic analysis alone, without any optimization.
# This *MAY* match temporal structure, pitch, and stationary formants
m1 = matchPars(target = target,
               samplingRate = 16000,
               maxIter = 0,  # no optimization, only acoustic analysis
               verbose = TRUE)
cand1 = do.call(soundgen, c(m1$pars, list(
  temperature = 0.001, play = TRUE, plot = TRUE)))

# Try to improve the match by optimizing rolloff
# (this may take a few minutes to run, and the results may vary)
m2 = matchPars(target = target,
               samplingRate = 16000,
               pars = 'rolloff',
               maxIter = 100,
               verbose = TRUE)
# rolloff should be moving from default (-12) to target (-20):
lapply(m2$history, function(x) x$pars$rolloff)
cand2 = do.call(soundgen, c(m2$pars, list(play = TRUE, plot = TRUE)))
}

Run the code above in your browser using DataLab