Learn R Programming

soundgen (version 3.0.0)

transplantFormants: Transplant formants

Description

Takes the general spectral envelope of one sound (donor) and "transplants" it onto another sound (recipient). For biological sounds like speech or animal vocalizations, this has the effect of replacing the formants in the recipient sound while preserving the original intonation and (to some extent) voice quality. Note that the amount of spectral smoothing (specified with freqWindow) is a crucial parameter: too little smoothing, and noise between harmonics will be amplified, creating artifacts; too much, and formants may be missed. The default is to set freqWindow to the estimated median pitch, but this is time-consuming and error-prone, so set it to a reasonable value manually if possible; if pitch detection fails, freqWindow defaults to 400 Hz. Also ensure that both sounds have the same sampling rate. You may want to fade the output a little (a very short linear fade-in/out is applied internally).

Usage

transplantFormants(
  donor,
  recipient,
  samplingRate = NULL,
  freqWindow = NULL,
  specEnvMethod = c("cepstral", "gauss", "movavg", "peak"),
  dynamicRange = 80,
  windowLength = 50,
  step = NULL,
  overlap = 75,
  wn = "gaussian",
  normalize = c("orig", "max", "none")
)

Value

The filtered waveform as a numeric vector with the original sampling rate, on a scale determined by the normalize argument and with the same duration as recipient.

Arguments

donor

either the sound that provides the formants (vector, Wave, or file) or the desired spectral filter (matrix) as returned by getFormantFilter or spectrogram - linear amplitude, frequency in rows, time in columns

recipient

the sound that receives the formants (vector, Wave, or file)

samplingRate

sampling rate (Hz) of both donor and recipient, which must match

freqWindow

the width of spectral smoothing window, Hz: a single positive number, roughly the expected spacing between harmonics. Defaults to the median pitch of the donor (or of the recipient if donor is a filter matrix); if pitch detection fails, defaults to 400 Hz with a message

specEnvMethod

the method of extracting a smoothed spectral envelope: "cepstral" = Gaussian liftering of the real cepstrum (default), "gauss" = Gaussian blur of the log spectrum, "movavg" = moving average of the log spectrum, "peak" = moving maximum (upper envelope). See getSpecEnv for details

dynamicRange

regions under -dynamicRange dB are treated as silent

windowLength

length of the analysis window, ms

step

step between successive windows, ms; if provided, overrides overlap; because digital audio is sampled at discrete time intervals of 1/samplingRate, the actual step and thus the time stamps of STFT frames may be slightly different - e.g., 24.98866 instead of 25.0 ms

overlap

overlap between successive windows, %

wn

window type accepted by winFun: character string or function

normalize

"orig" = same as donor / recipient (default), "max" = max possible amplitude of the donor given its scale (or of the recipient if donor is a filter matrix), "none" = no normalization

Details

Algorithm: makes spectrograms of both sounds, flattens the recipient spectrogram by dividing out its smoothed spectral envelope (obtained with getSpecEnv), smooths the donor spectrogram (or interpolates the supplied filter matrix) with getSpecEnv, multiplies the spectrograms, and transforms back into time domain with inverse STFT. To avoid amplifying noise, spectral bins more than dynamicRange dB below the peak of their frame are left untouched, and the original amplitude of each recipient frame is preserved. Anything more than dynamicRange dB below the global maximum is then zeroed out.

See Also

transplantEnv getFormantFilter addFormants getSpecEnv shiftFormants shiftPitch

Examples

Run this code
rec = rnorm(5000)  # white noise
donor = soundgen()  # voiced /a/
whisper = transplantFormants(donor = donor, recipient = rec,
  samplingRate = 16000, freqWindow = 300)  # whispered /a/
# playme(whisper)
meanSpectrum(whisper, 16000)

if (FALSE) {
# Objective: take formants from one sound and apply them to another
s_orig = soundgen(pitch = 100, formants = 'ai')

recipient = soundgen(
  sylLen = 1200,
  pitch = c(100, 300, 250, 200),
  vibratoFreq = 9, vibratoDep = 1,
  formants = NULL,
  addSilence = 180,
  samplingRate = 16000,  # same as donor
  invalidArgAction = 'ignore')  # force to keep the low samplingRate
playme(recipient, 16000)
spectrogram(recipient, 16000)

s1 = transplantFormants(
  donor = s_orig,
  recipient = recipient,
  samplingRate = 16000)
playme(s1, 16000)
spectrogram(s1, 16000)

# The spectral envelope of s1 will be similar to that of the original on a
# frequency scale determined by freqWindow. Compare the spectra:
par(mfrow = c(1, 2))
meanSpectrum(s_orig, 16000, yScale = 'max0', ylim = c(-50, 0), main = 'Donor')
meanSpectrum(s1, 16000, yScale = 'max0', ylim = c(-50, 0),
             main = 'Processed recipient')
par(mfrow = c(1, 1))

# if needed, transplant amplitude envelopes as well:
s2 = transplantEnv(donor = s_orig, recipient = s1,
                   samplingRateR = 16000, samplingRateD = 16000,
                   windowLength = 10)
playme(s2, 16000)
spectrogram(s2, 16000)
}

Run the code above in your browser using DataLab