Learn R Programming

tclust (version 2.2-3)

FowlkesMallowsIndex: Computes the Fowlkes and Mallows index

Description

Fowlkes-Mallows index is an external evaluation method that is used to determine the similarity between two clusterings (clusters obtained after a clustering algorithm). This measure of similarity could be either between two hierarchical clusterings or a clustering and a benchmark classification. A higher the value for the Fowlkes-Mallows index indicates a greater similarity between the clusters and the benchmark classifications. This index can be used to compare either two cluster label sets or a cluster label set with a true label set. The formula of the adjusted Fowlkes-Mallows index (ABk) is given in the details.

Usage

FowlkesMallowsIndex(c1, c2 = NULL, noisecluster = NULL)

Value

A list containing the following components:

ABK

Adjusted Fowlkes and Mallows index. A number between -1 and 1. The adjusted Fowlkes and Mallows index is the corrected-for-chance version of the Fowlkes and Mallows index.

BK

Value of the Fowlkes and Mallows index. A number between 0 and 1.

EBk

Expected value of the index computed under the null hypothesis of no-relation.

VarBk

Variance of the Fowlkes and Mallows index. Variance of the index computed under the null hypothesis of no-relation.

Arguments

c1

Labels of the first partition or contingency table. A numeric or character vector containing the class labels of the first partition or a 2-dimensional numeric matrix which contains the cross-tabulation of cluster assignments.

c2

Labels of the second partition. A numeric or character vector containing the class labels of the second partition. The length of vector c2 must be equal to the length of vector c1. This second input is required only if c1 is not a 2-dimensional numeric matrix.

noisecluster

Label or number associated to the noise class or noise level. Number or character label which denotes the points which do not belong to any cluster. These points are not takern into account for the computation of the Fowlkes and Mallows index

Details

The formula of the adjusted Fowlkes-Mallows index (ABk) is as follows:

$$ABk=\frac{Bk-Expected~value~of~Bk}{Max~Index - Expected~value~of~Bk}$$

References

Fowlkes, E.B. and Mallows, C.L. (1983), A Method for Comparing Two Hierarchical Clusterings, Journal of the American Statistical Association, Vol. 78, pp. 553--569.

See Also

randIndex

Examples

Run this code

 ##  1. FowlkesMallowsIndex (adjusted) with the two vectors as input
 c <- matrix(c(1, 1, 1, 2, 2, 1,2, 2, 2, 2, 2, 3, 3, 3, 3, 3, 3, 3, 3, 3), 
         ncol=2, byrow=TRUE)

 ##  c1 - numeric vector containing the labels of the first partition
 c1 <- c[, 1]

 ##  c2 - numeric vector containing the labels of the second partition
 c2 <- c[, 2]

 (FM <- FowlkesMallowsIndex(c1, c2))

 ##  2. FM index (adjusted) with the contingency table as input.
 T <- matrix(c(1, 1, 0, 1, 2, 1, 0, 0, 4), ncol=3, byrow=TRUE)
 (FM <- FowlkesMallowsIndex(T))


 ##  3. Compare FM (unadjusted) for iris data (true classification against 
 ##      tclust classification).

 ##  First partition c1 is the true partition
 c1 <- iris$Species

 ##  Second partition c2 is the output of tclust clustering procedure
 out <- tclust(iris[, 1:4], k=3, alpha=0, restr.fact=100)
 c2<- out$cluster

 (FM <- FowlkesMallowsIndex(c1, c2))

 ##  4. Compare FM index (unadjusted) for iris data (exclude unassigned units from tclust).

 ##  First partition c1 is the true partition
 c1 <- iris$Species

 ##  Second partition c2 is the output of tclust clustering procedure
 out <- tclust(iris[, 1:4], k=3, alpha=0.1, restr.fact=100)
 c2<- out$cluster

 ##  Units inside c2 which contain number 0 are referred to trimmed observations
 noisecluster <- 0
 (FM <- FowlkesMallowsIndex(c1, c2, noisecluster=noisecluster))

Run the code above in your browser using DataLab