Canonical correlation

Canonical correlation analysis between two variable sets

Canonical correlation asks a symmetric question: given two blocks of measured variables, how strongly do they relate, axis by axis? Magic Stat runs the classic cancor analysis from two spreadsheets and reports the correlations, the loadings and the significance tests.

The honest part

What canonical correlation is, and how to read it honestly

Canonical correlation takes two sets of continuous variables — call them X and Y — and finds the pair of linear combinations, one from each set, that correlate most strongly. It then repeats the process on the residuals, giving a sequence of canonical correlations, each orthogonal to the last.

This is a symmetric method. There is no response and no predictor: it describes how two blocks of measurements relate. That framing matters, because as soon as one block is meant to explain the other, a regression or a constrained ordination such as RDA is usually the more interpretable tool.

It is also a method with a deserved reputation for awkwardness. The raw coefficients are sensitive to collinearity, their signs can flip under trivial changes to the data, and with many variables relative to observations the solution is unstable. Read the canonical correlations and the structure loadings — the correlation of each original variable with the canonical scores — and treat the raw coefficients with caution.

The implementation here is the classic one: mean-centred X and Y, the generalised eigenvalue problem solved through the same SVD approach the reference implementation uses, raw coefficients, structure loadings, and the Wilks' lambda with Bartlett chi-square tests per root.

What it does

What Magic Stat gives you for canonical correlation

How it works

Two blocks to canonical correlations

  1. Load the data. both variable sets as columns of the same spreadsheet, one row per observation.
  2. Pick Canonical Corr. Statistics → Multivariate → Canonical Corr.
  3. Select X and Y. at least two numeric variables in each set, with no variable shared between the two.
  4. Run it. Magic Stat centres both sets, solves for the canonical roots and runs the Wilks and Bartlett tests.
  5. Read and export. correlations, coefficients, loadings and test statistics in the results tables; the summary goes into the report.
Options

The settings, in the dialog

Two blocks of continuous variables: at least two in X and at least two in Y, the same number of rows, and more observations than variables in each block.

No transformation step: the method centres X and Y exactly as the classic implementation does, so the transformation menu does not apply here. Missing values are handled by dropping incomplete rows consistently across both sets.

Frequently asked

Canonical correlation questions, answered honestly

How many canonical correlations will I get?

As many as the smaller of the two variable sets. The first is always the largest; each later root is uncorrelated with the ones before it and explains what is left.

What does Wilks' lambda test?

It tests the set of roots from a given point on: the first p-value tests all roots together, the next tests all but the first, and so on. A small p does not tell you which variables drive the relationship — the loadings do that.

The signs of my coefficients vary — is something broken?

No. Canonical variates are defined up to a sign flip, so coefficients and loadings can appear with either sign, and two runs on the same data can legitimately differ in sign. Interpret magnitudes and loadings, not signed coefficients.

Canonical correlation or CCA?

The names overlap badly. Canonical correlation (this page) relates two blocks of continuous variables symmetrically. CCA on our other page is canonical correspondence analysis, a constrained ordination of a count table. Choose by the data, not by the acronym.

What does it cost?

US$990 per year, with a free 48-hour trial and no card required. If you already work in R, stats::cancor and the CCP package reproduce this analysis and offer the same tests.

Canonical correspondence analysis → Redundancy analysis → Correlation analysis →