A simulated bivariate data set of size n = 100 and
dimension p = 2 from a k = 3 component mixture obtained
by applying the MixSim method of Maitra and Melnykov (2010),
as extended by Riani et al. (2015) and incorporated into the FSDA toolbox
of Matlab (Riani et al., 2012). The data set has been generated by imposing
an average cluster overlap (defined as a sum of pairwise misclassification
probabilities) equal to 0.04 and a maximum eigenvalue ratio for the scatters
matrices equal to 5.
data(mixsym)A data frame with 10 rows and 3 variables: 2 numerical and one categorical - the thrue cluster. The variables are as follows:
X1: first numerical variable
X2: second numerical variable
class: the true cluster
Simulated bivariate data set
Maitra, R. and Melnykov, V. (2010). Simulating data to study performance of finite mixture modeling and clustering algorithms. J. Comput. Graph. Stat., 19:354– 376.
Riani, M., Cerioli, A., Perrotta, D., and Torti, F. (2015). Simulating mixtures of multivariate data with fixed cluster overlap in FSDA library. Adv. Data Anal. Classif., 9:2015.
Riani, M., Perrotta, D., and Torti, F. (2012). FSDA: a matlab toolbox for robust analysis and interactive data exploration,. Chemometr. Intell. Lab. Syst., 116:17–32.