Non-parametric tests

Rank-based tests without the normality assumption

Non-parametric tests work from the ranks of the data rather than from the values themselves, so they do not assume a normal distribution. They are the right answer for ordinal data, strong skew and outliers — and the wrong reflex when the data are perfectly well behaved and a t-test would use more of the information.

The honest part

What rank-based tests assume, and what they cost

A non-parametric test replaces the observed values with their ranks and asks whether the ranks differ between groups. Because it does not depend on the shape of the distribution, it is robust to skew, to outliers and to ordinal scales where the distances between categories are not meaningful. The price is power: when the data really are normal, the rank-based test discards the information carried by the distances and can need a larger sample to detect the same effect.

The choice within the family depends on the design and on what you are willing to assume. Mann-Whitney U and Wilcoxon Rank-Sum (the same test under two names) compare two independent groups; Brunner-Munzel is the safer choice when the two distributions differ in shape as well as location. Kruskal-Wallis extends the rank comparison to three or more groups, and Mood's median test compares medians rather than ranks, at the cost of some power. The one-sample Wilcoxon signed-rank replaces the one-sample t-test, and Friedman replaces the repeated measures ANOVA.

A permutation test is available as a third route: it builds the null distribution by reshuffling the data, so it does not assume a distributional form either. It answers a slightly different question — a difference in means under exchangeability — which is worth knowing when you choose it.

What it does

What Magic Stat gives you for non-parametric tests

How it works

Pick the design, get the rank-based answer

  1. Load your data. the response in one column and the grouping factor in another, or the paired/repeated columns for the within-subjects route.
  2. Open Statistics → Means Comparison Tests. and choose the scope: Single Mean, Two Means or Multiple Means.
  3. Choose the non-parametric test. Mann-Whitney U, Wilcoxon Rank-Sum or Brunner-Munzel for two groups; Kruskal-Wallis or Mood's median test for three or more; the signed-rank test for one sample.
  4. Add a permutation test if you want one. tick the permutation option to get a p-value built by reshuffling the data.
  5. Read and export. the rank-based statistic, the effect size and the report, to .docx/.html/.md.
Options

The tests, in the dialog

Two independent groups: Mann-Whitney U · Wilcoxon Rank-Sum · Brunner-Munzel.

Three or more groups: Kruskal-Wallis · Mood's median test.

One sample: Wilcoxon signed-rank.

Repeated measures: Friedman (three or more conditions); Wilcoxon signed-rank (two conditions).

Permutation test: available for the two-sample and one-way routes, with the number of permutations reported.

Frequently asked

Non-parametric questions, answered honestly

When should I use a non-parametric test?

When the data are ordinal, strongly skewed, or dominated by outliers, or when the sample is too small for the normality assumption to be credible. If the data are roughly normal, a parametric test uses more of the information and is usually the better choice — using a rank test by reflex discards power you already have.

Should I transform the data instead?

A transformation can make a parametric test appropriate, but it changes the scale and therefore the meaning of the estimate, and it is not a cure for everything. A rank-based test is a cleaner statement about the data as measured. The app offers both routes so the decision is explicit.

Mann-Whitney or Brunner-Munzel?

Mann-Whitney compares locations when the two distributions have a similar shape. Brunner-Munzel is the safer choice when the distributions differ in shape as well, because it does not assume equal variances in the rank sense.

Can I get a p-value by permutation?

Yes. The permutation test builds the null distribution by reshuffling the group labels and is reported next to the formula-based result, which is a useful check when the sample is small or the distribution is uncertain.

Does it still report an effect size?

Yes. The multi-group case reports eta-squared, omega-squared and epsilon-squared; the two-group case reports Cohen's d with a confidence interval built by bootstrap or normal approximation. An effect size belongs with any test, not only with the parametric ones.

t-test → ANOVA → Post-hoc tests →