Learn R Programming

soundgen (version 3.0.0)

segment: Segment a sound

Description

Finds syllables and bursts / beats separated by background noise in long recordings (up to 1-2 hours of audio per file). Syllables are defined as continuous segments that differ from background noise based on amplitude and/or spectral contrast. Bursts are defined as local maxima in the detection contour that are high enough relative to the surrounding region. A note on long recordings: an hour of audio takes ~1 min to process, but watch your memory usage and perhaps decrease maxDur to avoid running out of RAM. This has little effect on processing speed, but dramatically cuts memory requirements. Another advantage of using shorter maxDur is that signal-noise separation may improve because noise amounts and profiles can be estimated in each chunk, adjusting for variable recording conditions over time. The downside is that small segments at chunk boundaries may be missed.

Usage

segment(
  x,
  samplingRate = NULL,
  scale = NULL,
  from = NULL,
  to = NULL,
  shortestSyl = 40,
  shortestPause = 40,
  input = c("mel", "env", "spec"),
  propNoise = NULL,
  SNR = NULL,
  specDiffMeasure = c("cosine", "logSpecDist", "specDiff"),
  amplWeight = c("add", "gate", "scale", "none"),
  noiseLevelStabWeight = c(1, 0.25),
  windowLength = 40,
  step = NULL,
  overlap = 80,
  reverb_pars = list(reverbDelay = 70, reverbSpread = 130, reverbLevel = -35,
    reverbDensity = 50, echoLevel = -Inf),
  interburst = NULL,
  peakToTrough = NULL,
  summaryFun = c("median", "sd"),
  maxDur = 120,
  ptvStep = NULL,
  ptvTime = 0.5,
  ptvFreq = 1,
  reportEvery = NULL,
  cores = 1,
  plot = FALSE,
  savePlots = FALSE,
  embed = FALSE,
  saveAudio = FALSE,
  addSilence = 50,
  main = NULL,
  xlab = "",
  ylab = NULL,
  showLegend = FALSE,
  width = 900,
  height = 500,
  units = "px",
  res = NA,
  maxPoints = c(1e+05, 5e+05),
  specPlot = list(colorTheme = "bw"),
  contourPlot = list(lty = 1, lwd = 2, col = "green"),
  sylPlot = list(lty = 1, lwd = 2, col = "blue"),
  burstPlot = list(pch = 8, cex = 3, col = "red"),
  ...
)

Value

A list with the following components:

syllables

data frame with columns syllable, start, end, pauseLen, sylLen, sylRate, and ptv. Times are in ms. start and end correspond to syllable edges; pauseLen is the time between the end of the previous syllable and the start of the current syllable. $ptv gives the proportion of time vocalizing for each syllable-pause pair.

bursts

data frame with columns time, ampl, and interburst. time and interburst are in ms

summary

data frame summarizing temporal descriptives per file, including global PTV = sum of all syllable durations / total duration of audio. NULL if summaryFun = NULL

ptv

contour of instantaneous proportion of time vocalizing, with columns time (ms), on (1 if this frame is part of a syllable, 0 otherwise), ptv_lowpass, and ptv_conv. If PTV cannot be computed (e.g., if there are <2 syllables), this may be NA or 0

If more than one file is analyzed, syllables, bursts, and

ptv are lists with one element per file, while summary remains a single data frame.

Arguments

x

path to a folder, one or more wav or mp3 files c('file1.wav', 'file2.mp3'), Wave object, numeric vector, or a list of Wave objects or numeric vectors

samplingRate

sampling rate of x (only needed if x is a numeric vector)

scale

maximum possible amplitude of input, used to normalize the input vector (only needed if x is a numeric vector)

from, to

if NULL (default), analyzes the whole sound, otherwise from...to (s)

shortestSyl

minimum acceptable length of syllables, ms

shortestPause

minimum acceptable break between syllables, ms (syllables separated by shorter pauses are merged)

input

the contour used to search for syllables: 'env' = smoothed RMS amplitude envelope; 'spec' = power spectrum returned by tuneR::melfcc; 'mel' = mel-filterbank amplitude spectrum from tuneR::melfcc

propNoise

the proportion of analysis frames assumed to represent background noise, 0 to 1. If NULL, this proportion is estimated automatically in each chunk. If 0, correction for background noise is skipped, in which case SNR must be specified

SNR

expected signal-to-noise ratio (dB above noise), which determines the threshold for syllable detection. If NULL, SNR is estimated automatically in each chunk unless propNoise = 0. The meaning of "dB" here is approximate because the detection contour may not be sound intensity, depending on input

specDiffMeasure

similarity measure used to compare each spectrum with the estimated noise spectrum (ignored when input = 'env'):

cosine

cosine distance without centering

logSpecDist

root-mean-square log-spectral distance

specDiff

mean absolute log-spectral difference

amplWeight

amplitude weighting applied to the spectral contrast (ignored when input = 'env'):

add

combines spectral distinctiveness with amplitude elevation; the default behavior in soundgen v2.x when specDiffMeasure = 'cosine'

gate

amplitude-gated contrast: spectral contrast contributes mainly when the frame is also above the estimated noise amplitude

scale

spectral contrast scaled to the dynamic range of the amplitude contour

none

unnormalized spectral contrast, mainly useful for diagnostics

noiseLevelStabWeight

a vector of length 2 specifying the relative weights of the overall signal level vs. stability (time derivative) when attempting to automatically locate the regions that represent noise. Increasing the weight of stability prioritizes sudden changes as marking the beginning and end of a syllable

windowLength

length of the analysis window, ms

step

step between successive windows, ms; if provided, overrides overlap; because digital audio is sampled at discrete time intervals of 1/samplingRate, the actual step and thus the time stamps of STFT frames may be slightly different - e.g., 24.98866 instead of 25.0 ms

overlap

overlap between successive windows, %

reverb_pars

parameters passed on to reverb to attempt to cancel the effects of reverberation or echo, which otherwise tend to merge short and loud segments like rapid barks. This is particularly useful if you have a recording of an impulse response in the same environment to pass on to reverb. The default disables echo by setting echoLevel = -Inf. Set to NULL to use a static detection threshold without reverb compensation

interburst

minimum time between two consecutive bursts, ms. Also determines the analysis window used for detecting burst peaks. Defaults to the median detected sylLen + pauseLen in each chunk, falling back to shortestSyl if no syllables are detected

peakToTrough

to qualify as a burst, a local maximum has to be at least peakToTrough dB above the surrounding contour. Defaults to SNR + 3 dB in each chunk when peakToTrough = NULL

summaryFun

functions used to summarize each acoustic characteristic; see analyze

maxDur

long files are split into chunks maxDur s in duration to avoid running out of RAM. The outputs for all fragments are glued together, but plotting is switched off for chunked files. Note that the noise profile is estimated in each chunk separately, so set it low if the background noise is highly variable

ptvStep, ptvTime, ptvFreq

the instantaneous proportion of time vocalizing (PTV) is calculated by producing a binary (sound on/off) contour with a step of ptvStep ms and convolving it with a half-Gaussian filter with SD = ptvTime seconds ($ptv_conv), and by low-pass filtering it over ptvFreq Hz ($ptv_lowpass). ptvStep defaults to the same value as step

reportEvery

when processing multiple inputs, report estimated time left every reportEvery iterations (NULL = default, NA = don't report); see reportTime

cores

number of cores for parallel processing

plot

if TRUE, produces a plot of the results

savePlots

if TRUE, creates a subdirectory in the input directory (if input is a file or folder) or in the working directory (if input is a vector etc), named after the function (eg "spectrogram/"). All plots and audio files (if any) are saved in this new directory. If there are multiple inputs, an html notebook is also created for easy viewing and listening

embed

if TRUE and savePlots is set and there are multiple inputs, all saved images and audio (if any) are embedded in the exported html notebook for easy sharing; if FALSE, the html file links to separate images and audio files (but separate files are still saved). NB: for this to work, package "base64enc" must be installed

saveAudio

if TRUE, saves the extracted syllables in a "/segment" subdirectory created in the input directory (if input is a file or folder) or in the working directory

addSilence

if syllables are saved as separate audio files, they are padded with this much silence before and after, ms

xlab, ylab, main

main plotting parameters

showLegend

if TRUE, shows a legend for thresholds

width, height, units, res

graphical parameters for saving plots passed to png

maxPoints

maximum number of points to plot for the waveform and the segmentation contour. Longer contours are downsampled to avoid slow plotting

specPlot

a list of graphical parameters for displaying the spectrogram (if input = 'spec' or 'mel'); set to NULL to hide the spectrogram

contourPlot

a list of graphical parameters for displaying the signal contour used to detect syllables

sylPlot

a list of graphical parameters for displaying the syllable threshold

burstPlot

a list of graphical parameters for displaying the bursts

...

other graphical parameters passed to graphics::plot

Details

Algorithm: the sound is analyzed in chunks of at most maxDur seconds. In each chunk, the quietest and most stable regions are located, and a noise threshold is derived either from a user-specified proportion of noise (propNoise) or, if propNoise = NULL, automatically from the distribution of a weighted product of amplitude and stability. The detection contour is then compared against the estimated noise. If input = 'env', the contour is a smoothed log RMS amplitude envelope. If input = 'spec' or 'mel', the contour is computed by comparing the spectrum of each frame with the estimated noise spectrum using specDiffMeasure and amplWeight.

Syllables are detected as continuous regions of the contour that exceed the noise threshold by approximately SNR dB. Pauses shorter than shortestPause are merged. Syllable start and end times correspond to the edges of envelope bins, and pauseLen is the time between the end of the previous syllable and the start of the current syllable.

Bursts are detected as local maxima of the contour above the syllable detection threshold. The minimum spacing/scale for burst detection is controlled by interburst, and the required peak prominence is controlled by peakToTrough.

See Also

analyze ssm segment_ann

Examples

Run this code
sound = soundgen(nSyl = 4, sylLen = 100, pauseLen = 70,
                 attackLen = 20, amplGlobal = c(0, -20),
                 pitch = c(368, 284), temperature = .001)
# add noise so SNR decreases from 20 to 0 dB from syl1 to syl4
sound = sound + runif(length(sound), -10 ^ (-20 / 20), 10 ^ (-20 / 20))
# osc(sound, samplingRate = 16000, dB = TRUE)
# spectrogram(sound, samplingRate = 16000)
# playme(sound, samplingRate = 16000)

s = segment(sound, samplingRate = 16000, plot = TRUE)
str(s)

# customizing the plot
segment(sound, samplingRate = 16000, plot = TRUE,
        sylPlot = list(lty = 2, col = 'gray20'),
        burstPlot = list(pch = 16, col = 'blue'),
        specPlot = list(col = rev(heat.colors(50))),
        xlab = 'Some custom label', cex.lab = 1.2,
        showLegend = TRUE,
        main = 'My awesome plot')

# set SNR manually to control detection threshold
s = segment(sound, samplingRate = 16000, SNR = 1, plot = TRUE)

# simple intensity threshold (anything >5 dB is signal)
segment(sound, 16000, input = 'env', SNR = 5, plot = TRUE,
  # less smoothing gives more precise timing
  windowLength = 10, step = 5,
  # don't correct SNR based on estimated background noise
  propNoise = 0,
  # don't use dynamic thresholds to cancel reverb
  reverb_pars = NULL
)

# same with automatic threshold setting
segment(sound, 16000, input = 'env', plot = TRUE,
  windowLength = 10, step = 5, reverb_pars = NULL)

if (FALSE) {
# plot the PTV contour (proportion of time vocalizing)
plot(s$ptv$time, s$ptv$on, type = 'l', xlab = 'Time, ms',
  ylab = 'Prop. time voc.')
points(s$ptv$time, s$ptv$ptv_conv, type = 'l', col = 'blue')
points(s$ptv$time, s$ptv$ptv_lowpass, type = 'l', col = 'red')
s$summary$ptv; mean(s$ptv$ptv_conv); mean(s$ptv$ptv_lowpass) # similar

# different ways to calculate instantaneous PTV
s2 = segment(sound, 16000, ptvTime = 2, ptvFreq = 5)
s3 = segment(sound, 16000, ptvTime = 0.05, ptvFreq = 0.5)
plot(s2$ptv$time, s2$ptv$on, type = 'l', xlab = 'Time, ms',
  ylab = 'Prop. time voc.')
points(s2$ptv$time, s2$ptv$ptv_conv, type = 'l', col = 'blue')
points(s2$ptv$time, s2$ptv$ptv_lowpass, type = 'l', col = 'yellow')
points(s3$ptv$time, s3$ptv$ptv_conv, type = 'l', col = 'purple')
points(s3$ptv$time, s3$ptv$ptv_lowpass, type = 'l', col = 'orange')

# segment all files in a folder and save the segments as separate files
segment('~/Downloads/temp',
  saveAudio = TRUE, savePlots = TRUE)
}

Run the code above in your browser using DataLab