Spikes are detected based on a modified \(Z\) score calculated
from the differenced spectrum. The \(Z\) threshold used should be
adjusted to the characteristics of the input and desired sensitivity. The
lower the threshold the more stringent the test becomes, with shorter
spikes being detected.
The algorithms assume a consistent step size for the underlying
independent variable, e.g., wavelength or time, and should not be applied
if the data do not fulfil this assumption, at least approximately. As
find_spkikes() operates on a single vector, checking this remains
the responsibility of calling functions or methods such as
spikes() and despike().
The algorithm uses running differences to detect abrupt changes in value,
compared to an estimate of the baseline variation of the differences,
approximating a baseline \(Z\) from MAD and a baseline value from the
median differences. Currently, a single estimate of MAD is used but running
medians, when possible, as baseline. This comparison detects running
differences that are unusually large, in most cases signalling a transition
between values near the baseline and far from it, in both directions.
Transitions into- and out of spikes are distinguished based on the median
of the non-differenced values, as a descriptor of the data baseline. As for
the median of the differences, a running median is used when possible.
This function thus detects the start and end of each spike, and
distinguishes upward and downward spikes.
k is the width in number of observations of the window used for
running median smoothing to extract the baseline. A value several times the
width of the broader spike but narrow enough to track broader peaks needs
to be manually set in most cases.
With na.rm = TRUE, NA values are omitted before searching for
spikes and set to 0L in the returned vector.
If all spikes are guaranteed to be one observation-wide and either going up
or down from the baseline, it is possible to detect them based purely on
the z.threshold by passing height.threshold = NA and either
spike.direction = "up" or spike.direction = "down", which
ensures very fast computation.
Parameters of the algorithm need to be adjusted depending on the data, so
inspection of returned values is needed together with adjustment by trial
and error of suitable values for z.threshold,
height.threshold, and k.
The argument to parameter max.spike.width is used in a final stage
to discard spikes detected by the algorithms described above but considered
to be too wide (as a run of successive observations, i.e., detector pixels
in the case of array spectrometers).