Learn R Programming

soundgen (version 3.0.0)

addFormants: Add formants

Description

A spectral filter that either adds or removes formants from a sound - that is, amplifies or dampens certain frequency bands, as in human vowels. See soundgen and getFormantFilter for more information. With action = 'remove' this function can perform inverse filtering to remove formants and obtain raw glottal output, provided that you can specify the correct formant structure. Instead of formants, any arbitrary spectral filtering function can be applied using the formantFilter argument (e.g., for a low/high/bandpass filter).

Usage

addFormants(
  x,
  samplingRate = NULL,
  formants = NULL,
  formantFilter = NULL,
  action = c("add", "remove"),
  dB = NULL,
  specificity = 1,
  zFun = NULL,
  vocalTract = NA,
  formantDep = 1,
  formantDepStoch = 1,
  formantWidth = 1,
  formantCeiling = NULL,
  lipRad = 6,
  noseRad = 4,
  mouthOpenThres = 0,
  mouth = NA,
  temperature = 0.025,
  formDrift = 0.3,
  formDisp = 0.2,
  smoothing = list(interpol = "splineFC"),
  windowLength = 50,
  step = NULL,
  overlap = 75,
  wn = "gaussian",
  normalize = c("orig", "max", "none"),
  play = FALSE,
  saveAudio = FALSE,
  reportEvery = NULL,
  cores = 1,
  ...
)

Value

The filtered waveform as a numeric vector of the original length with the original sampling rate, or a list if there are multiple inputs.

Arguments

x

path to a folder, one or more wav or mp3 files c('file1.wav', 'file2.mp3'), Wave object, numeric vector, or a list of Wave objects or numeric vectors

samplingRate

sampling rate of x (only needed if x is a numeric vector)

formants

a vector of formant frequencies (assuming formants are static throughout the sound); a list of formant times, frequencies, amplitudes, and bandwidths; or a character string referring to default presets for speaker "M1" (implemented: "aoieu0"). NA or NULL means no formants, only lip radiation (but a schwa is generated if vocalTract is specified). Time stamps for formants and mouth can be specified in ms relative to sylLen or on a scale of [0, 1]. See getFormantFilter for more details

formantFilter

(optional): as an alternative to specifying formant frequencies, we can provide the exact filter - a vector of non-negative numbers specifying the amplitude in each frequency bin on a linear scale. A matrix specifying the filter for each STFT step with frequency bins in rows and STFT frames in columns is also accepted. The easiest way to create this matrix is to call getFormantFilter or to use the spectrum of a recorded sound

action

'add' = add formants to the sound (default), 'remove' = remove formants (inverse filtering)

dB

if NULL (default), the spectral envelope is applied on the original scale; otherwise, it is set to range up to 10^(dB / 20)

specificity

a way to sharpen or blur the spectral envelope (spectrum ^ specificity) : 1 = no change, >1 = sharper, <1 = blurred

zFun

(optional) an arbitrary function to apply to the spectrogram prior to iSTFT, where "z" is the spectrogram - a matrix of complex values (see examples)

vocalTract

the length of vocal tract, cm. Used for calculating formant dispersion (for adding extra formants) and formant transitions as the mouth opens and closes. If NULL or NA, the length is estimated based on specified formant frequencies, if any (anchor format)

formantDep

scale factor of formant amplitude (1 = no change relative to amplitudes in formants)

formantDepStoch

the amplitude of additional stochastic formants added above the highest specified formant, dB (only if temperature > 0)

formantWidth

scale factor of formant bandwidth (1 = no change)

formantCeiling

frequency to which stochastic formants are calculated to avoid losing energy in the upper part of the spectrum due to unmodeled resonances above the Nyquist frequencies, specified in multiples of Nyquist. If NULL (default), a faster theoretical correction is used (may fail for unusual sounds)

lipRad

the effect of lip radiation on source spectrum, dB/oct (the default of +6 dB/oct produces a high-frequency boost when the mouth is open)

noseRad

the effect of radiation through the nose on source spectrum, dB/oct (the alternative to lipRad when the mouth is closed)

mouthOpenThres

open the lips (switch from nose radiation to lip radiation) when the mouth is open >mouthOpenThres, 0 to 1

mouth

mouth opening (0 to 1, 0.5 = neutral, i.e. no modification) (anchor format)

temperature

hyperparameter for regulating the amount of stochasticity in sound generation

formDrift, formDisp

scaling factors for the effect of temperature on formant drift and dispersal, respectively

smoothing

a list of parameters passed to interpolate to control the interpolation and smoothing of contours drawn through anchors

windowLength

length of the analysis window, ms

step

step between successive windows, ms; if provided, overrides overlap; because digital audio is sampled at discrete time intervals of 1/samplingRate, the actual step and thus the time stamps of STFT frames may be slightly different - e.g., 24.98866 instead of 25.0 ms

overlap

overlap between successive windows, %

wn

window type accepted by winFun: character string or function

normalize

"orig" = same as input (default), "max" = maximum possible peak amplitude for the given input scale, "none" = no normalization

play

if TRUE, plays the output audio using the default player on your system. If a character string, it is passed to playme as the name of the player to use (e.g. 'aplay', 'play', 'vlc'). In case of errors, try setting another default player for playme

saveAudio

if TRUE, saves the processed audio in a subdirectory named after the function and created in the input directory (if input is a file or folder) or in the working directory

reportEvery

when processing multiple inputs, report estimated time left every reportEvery iterations (NULL = default, NA = don't report); see reportTime

cores

number of cores for parallel processing

...

extra parameters passed to zFun

Details

Algorithm: converts input from a time series (time domain) to a spectrogram (frequency domain) through short-time Fourier transform (STFT), multiplies by the spectral filter containing the specified formants, and transforms back to a time series via inverse STFT. This is a subroutine for voice synthesis in soundgen, but it can also be applied to a recording.

See Also

getFormantFilter transplantFormants soundgen

Examples

Run this code
sound = c(rep(0, 1000), rnorm(8000) * 2 - 1, rep(0, 1000))  # white noise
# NB: pad with silence to avoid artifacts if removing formants
# playme(sound)
# spectrogram(sound, samplingRate = 16000)

# add F1 = 900, F2 = 1300 Hz
sound_filtered = addFormants(sound, samplingRate = 16000,
                             formants = c(900, 1300))
# playme(sound_filtered)
# spectrogram(sound_filtered, samplingRate = 16000)

# ...and remove them again (assuming we know what the formants are)
sound_inverse_filt = addFormants(sound_filtered,
                                 samplingRate = 16000,
                                 formants = c(900, 1300),
                                 action = 'remove')
# playme(sound_inverse_filt)
# spectrogram(sound_inverse_filt, samplingRate = 16000)

if (FALSE) {
## Perform some user-defined manipulation of the spectrogram with zFun
# Ex.: noise removal - silence all bins 50 dB below the max value
s_noisy = soundgen(sylLen = 200, addSilence = 0,
                   noise = list(time = c(-100, 300), value = -20))
spectrogram(s_noisy, 16000)
# playme(s_noisy)
zFun = function(z, cutoff = -50) {
  az = abs(z)
  thres = max(az) * 10 ^ (cutoff / 20)
  z[which(az < thres)] = 0
  return(z)
}
s_denoised = addFormants(s_noisy, samplingRate = 16000,
                         formants = NA, zFun = zFun, cutoff = -40)
spectrogram(s_denoised, 16000)
# playme(s_denoised)

# If neither formants nor formantFilter are defined, only lipRad has an effect
# For ex., we can boost low frequencies by 6 dB/oct
noise = rnorm(8000)
noise1 = addFormants(noise, 16000, lipRad = -6)
meanSpectrum(noise1, 16000, yScale = 'max0')

# Arbitrary spectra can be defined with formantFilter. For ex., we can
# have a flat spectrum up to 2 kHz (Nyquist / 4) and -3 dB/kHz above:
freqs = seq(0, 16000 / 2, length.out = 100)
n = length(freqs)
idx = (n / 4):n
sp_dB = c(rep(0, n / 4 - 1), (freqs[idx] - freqs[idx[1]]) / 1000 * (-3))
plot(freqs, sp_dB, type = 'b')
noise2 = addFormants(noise, 16000, lipRad = 0, formantFilter = 10 ^ (sp_dB / 20))
meanSpectrum(noise2, 16000, yScale = 'max0')

## Use the spectral envelope of another recording
# (NB: this can also be achieved with a single call to transplantFormants)
sound_orig = soundgen(sylLen = 300, formants = 'a', addSilence = 5)
samplingRate = 16000
# playme(sound_orig, samplingRate)

# get a few pitch anchors to reproduce the original intonation
pitch = analyze(sound_orig, samplingRate = samplingRate,
  pitchMethod = c('autocor', 'dom'))$detailed$pitch
pitch = pitch[!is.na(pitch)]

# extract a frequency-smoothed version of the original spectrogram
# to use as filter
specEnv_orig = spectrogram(sound_orig, blur = c(300, 50),
 samplingRate = samplingRate, output = 'original', plot = TRUE)

# Synthesize source only, with flat spectrum
sound_unfilt = soundgen(sylLen = 2500, pitch = pitch,
  rolloff = 0, rolloffOct = 0,
  temperature = 0, formants = NULL, lipRad = 0,
  samplingRate = samplingRate,
  invalidArgAction = 'ignore')  # prevent soundgen from increasing samplingRate
# playme(sound_unfilt, samplingRate)
# meanSpectrum(sound_unfilt, samplingRate, yScale = 'max0')  # ~flat

# Force spectral envelope to the shape of target
sound_filt = addFormants(sound_unfilt, formants = NULL,
  formantFilter = specEnv_orig, samplingRate = samplingRate)
# playme(sound_filt, samplingRate)  # playme(sound_orig, samplingRate)
# spectrogram(sound_filt, samplingRate)  # spectrogram(sound_orig, samplingRate)

# The spectral envelope is now similar to the original recording. Compare:
par(mfrow = c(1, 2))
meanSpectrum(sound_orig, samplingRate, yScale = 'max0', alim = c(-50, 20))
meanSpectrum(sound_filt, samplingRate, yScale = 'max0', alim = c(-50, 20))
par(mfrow = c(1, 1))
}

Run the code above in your browser using DataLab