A less precise, but very quick method of pitch tracking based on measuring
zero-crossing rate in bandpass-filtered audio. Recommended for processing
long recordings with typical pitch values well below the first formant
frequency, such as speech. Calling this function is considerably faster than
using the same pitch-tracking method in analyze. Note that,
unlike analyze(), it returns the times of individual zero crossings
(hopefully corresponding to glottal cycles) instead of pitch values at fixed
time intervals.
getPitchZc(
x,
samplingRate = NULL,
scale = NULL,
from = NULL,
to = NULL,
pitchFloor = 50,
pitchCeiling = 400,
zcThres = 0.1,
zcWin = 5,
silence = 0.04,
envWin = 5,
certMethod = c("variab", "autocor"),
summaryFun = c("mean", "sd"),
reportEvery = NULL
)A list with descriptives per file (@summary) and per frame (@detailed), including
pitch calculated from the time between consecutive zero crossings
certainty in each pitch candidate calculated from local pitch stability, 0 to 1
path to a folder, one or more wav or mp3 files c('file1.wav', 'file2.mp3'), Wave object, numeric vector, or a list of Wave objects or numeric vectors
sampling rate of x (only needed if x is a
numeric vector)
maximum possible amplitude of input, used to normalize the input
vector (only needed if x is a numeric vector)
if NULL (default), analyzes the whole sound, otherwise from...to (s)
absolute bounds for pitch candidates (Hz)
pitch candidates with certainty below this value are treated
as noise and set to NA (0 = anything goes, 1 = pitch must be perfectly
stable over zcWin)
certainty in pitch candidates depends on how stable pitch is
over zcWin glottal cycles (odd integer > 3)
minimum root mean square (RMS) amplitude, below which pitch candidates are set to NA (NULL = don't consider RMS amplitude)
window length for calculating RMS envelope, ms
method of calculating pitch certainty: 'variab' = variability of pitch estimates per zc over window (default); 'autocor' = autocorrelation of pitch estimates per zc over window (a measure of curve smoothness)
functions used to summarize each acoustic characteristic,
eg c('mean', 'sd'); user-defined functions are fine (see examples);
NAs are omitted automatically for mean/median/sd/min/max/range/sum,
otherwise take care of NAs yourself
when processing multiple inputs, report estimated time
left every reportEvery iterations (NULL = default, NA = don't
report); see reportTime
Algorithm: the audio is bandpass-filtered from pitchFloor to
pitchCeiling, and the timing of all zero crossings is saved. This is
not enough, however, because aperiodic sounds like white noise also have
plenty of zero crossings. Accordingly, an attempt is made to detect voiced
segments (or steady musical tones, etc.) by looking for stable regions, with
several zero-crossings at relatively regular intervals (see parameters
zcThres and zcWin). Very quiet parts of audio are also treated
as not having a pitch.
analyze
data(speechEx, package = 'soundgen')
# spectrogram(speechEx)
zc = soundgen:::getPitchZc(speechEx, pitchCeiling = 250)
plot(zc$detailed[, c('time', 'pitch')], type = 'b')
spectrogram(speechEx, extraContour = zc$detailed$pitch, ylim = c(0, 2))
if (FALSE) {
# process all files in a folder
zc = soundgen:::getPitchZc('~/Downloads/temp')
zc$summary
}
Run the code above in your browser using DataLab