Cluster analysis

Cluster analysis without writing a script

Cluster analysis finds groups in data when no grouping is given. Magic Stat runs both families: hierarchical clustering with a dendrogram and a k-cluster cut, and k-means with a fixed seed and a silhouette diagnostic.

The honest part

Two families of clustering, and when each is the right one

Hierarchical clustering builds a tree: at every step it merges the two closest clusters, using a linkage rule such as Ward, complete, average or single linkage, and the result is a dendrogram you cut at k clusters. It assumes nothing about the shape of the groups, but it is greedy — an early wrong merge is never undone.

K-means does the opposite: you fix k in advance, and the algorithm iterates assignments and centroids to minimise the total within-cluster sum of squares. It is fast and gives a tidy partition, but it assumes roughly spherical clusters of similar size, and it is sensitive to the starting points — which is why the seed is fixed here and the run is reproducible.

The linkage and distance choices in hierarchical clustering are not neutral. Ward linkage on Euclidean distances behaves differently from average linkage on Bray–Curtis, and a poor pairing can manufacture groups that are not there. The dialog makes both choices explicit.

Clustering is exploratory. It will return clusters even from structureless data, so the number of clusters and the stability of the solution deserve as much attention as the dendrogram. If you already have groups and want to know whether they differ, that is a different job — PERMANOVA or discriminant analysis.

What it does

What Magic Stat gives you for clustering

How it works

From a spreadsheet to a partition

  1. Load the data. observations in rows, numeric variables in columns; incomplete rows are removed with a warning.
  2. Pick Cluster. and choose hierarchical clustering or k-means.
  3. Set the options. for hierarchical, the linkage and the distance measure; for k-means, the number of clusters and the seed.
  4. Choose k. the number of clusters for the cut, or for k-means directly.
  5. Read and export. the assignment table, the silhouette and the plot; the result goes into the report.
Options

The settings, in the dialog

Hierarchical: linkage method (ward.D2 · complete · average · single · centroid · median · mcquitty) and distance measure (Bray–Curtis · Euclidean · Manhattan · Canberra), cut at k clusters.

K-means: number of clusters and a random seed (fixed by default for reproducibility); the Lloyd algorithm is used, matching the classic reference implementation. The variables are used as given, so standardise them beforehand if they are on different scales.

Frequently asked

Cluster analysis questions, answered honestly

How many clusters should I use?

There is no automatic answer, and the dialog will not pretend otherwise. Use the silhouette value, the shape of the dendrogram and the merge heights, and — most importantly — whether the clusters mean something in your field. Report the criterion you used.

Do I need to standardise the variables?

If they are on different scales, yes — otherwise a variable measured in large numbers dominates the distance. The dialog does not standardise for you, so scale the columns first if that is what your question needs.

Hierarchical or k-means?

Use hierarchical when you want a tree and do not want to fix k in advance, and k-means when you have a reason to fix k and want a fast, tidy partition. They often agree; when they do not, the disagreement is itself worth reporting.

Will it find clusters even in random data?

Yes — every clustering method partitions whatever you give it. A dendrogram is not evidence that groups exist. Checking stability and using an independent criterion is what turns a partition into a finding.

What does it cost?

US$990 per year with a free 48-hour trial and no card. If your analysis needs methods that are not here — fuzzy clustering, model-based clustering, or the very large-n optimisations — R's cluster and mclust packages are the right place to look.

PCA analysis → NMDS analysis → Discriminant analysis →