Learn R Programming

soundgen (version 3.0.0)

prosody: Prosody

Description

Exaggerates or flattens the intonation by performing a dynamic pitch shift, changing pitch excursion from its original median value without changing the formants. This is a particular case of pitch shifting, which is performed with shiftPitch. The result is likely to be improved if manually corrected pitch contours are provided. Depending on the nature of audio, the settings that control pitch shifting may also need to be fine-tuned with the shiftPitch_pars argument. Any NAs in pitch contour are treated as voiceless fragments, and their pitch is not modified. If the NAs are not actual silences / voiceless frames, they should be interpolated before passing the pitch contour to prosody().

Usage

prosody(
  x,
  samplingRate = NULL,
  multProsody,
  analyze_pars = list(),
  shiftPitch_pars = list(),
  pitchManual = NULL,
  reportEvery = NULL,
  cores = 1,
  play = FALSE,
  saveAudio = FALSE,
  plot = FALSE,
  savePlots = FALSE,
  embed = FALSE,
  width = 900,
  height = 500,
  units = "px",
  res = NA
)

Value

If the input is a single audio (file, Wave, or numeric vector), returns the processed waveform as a numeric vector with the original sampling rate and scale. If the input is a folder with several audio files, returns a list of processed waveforms, one for each file.

Arguments

x

path to a folder, one or more wav or mp3 files c('file1.wav', 'file2.mp3'), Wave object, numeric vector, or a list of Wave objects or numeric vectors

samplingRate

sampling rate of x (only needed if x is a numeric vector)

multProsody

multiplier of pitch excursion from median (on a logarithmic or musical scale): >1 = exaggerate intonation, 1 = no change, <1 = flatten, 0 = completely flat at the original median pitch; numeric vector or anchor format list(time = ..., value = ...)

analyze_pars

a list of parameters to pass to analyze (only needed if pitchManual is NULL - that is, if we attempt to track pitch automatically)

shiftPitch_pars

a list of parameters to pass to shiftPitch to fine-tune the pitch-shifting algorithm

pitchManual

manually corrected pitch contour. For a single sound, provide a numeric vector of any length. For multiple sounds, provide a dataframe with columns "file" and "pitch" (or path to a csv file) as returned by pitch_app, ideally with the same windowLength and step as in current call to analyze. A named list with pitch vectors per file is also accepted - e.g., as returned by pitch_app

reportEvery

when processing multiple inputs, report estimated time left every reportEvery iterations (NULL = default, NA = don't report); see reportTime

cores

number of cores for parallel processing

play

if TRUE, plays the output audio using the default player on your system. If a character string, it is passed to playme as the name of the player to use (e.g. 'aplay', 'play', 'vlc'). In case of errors, try setting another default player for playme

saveAudio

if TRUE, saves the processed audio in a subdirectory named after the function and created in the input directory (if input is a file or folder) or in the working directory

plot

if TRUE, produces a plot of the results

savePlots

if TRUE, creates a subdirectory in the input directory (if input is a file or folder) or in the working directory (if input is a vector etc), named after the function (eg "spectrogram/"). All plots and audio files (if any) are saved in this new directory. If there are multiple inputs, an html notebook is also created for easy viewing and listening

embed

if TRUE and savePlots is set and there are multiple inputs, all saved images and audio (if any) are embedded in the exported html notebook for easy sharing; if FALSE, the html file links to separate images and audio files (but separate files are still saved). NB: for this to work, package "base64enc" must be installed

width, height, units, res

graphical parameters for saving plots passed to png

See Also

shiftPitch

Examples

Run this code
s = soundgen(sylLen = 200, pitch = c(150, 220), addSilence = 50,
             plot = TRUE, yScale = 'log')
# playme(s)
s1 = prosody(s, 16000, multProsody = 2,
  analyze_pars = list(windowLength = 30, step = 15),
  shiftPitch_pars = list(windowLength = 20, step = 5, freqWindow = 300),
  plot = TRUE)
# playme(s1)
# spectrogram(s1, 16000, yScale = 'log')

if (FALSE) {
data('speechEx', package = 'soundgen')
samplingRate = [email protected]
spectrogram(speechEx, yScale = 'log', ylim = c(.05, 4))
# playme(speechEx)

# start with exaggerated prosody, then flat towards the end
speech1 = prosody(speechEx,
  multProsody = list(time = c(0, 1), value = c(1.5, 0.1)),
  analyze_pars = list(windowLength = 40, step = 10,
  pitchMethods = c('dom', 'autocor', 'cep', 'spec')),
  shiftPitch_pars = list(freqWindow = 400))
spectrogram(speech1, samplingRate, yScale = 'log')
# playme(speech1, samplingRate)

# process all audio files in a folder
s4 = prosody('~/Downloads/temp', multProsody = 2,
  savePlots = TRUE, saveAudio = TRUE)
str(s4)  # returns a list with audio (+ saves it to disk)
}

Run the code above in your browser using DataLab