A significant ANOVA says that some group means differ; it does not say which. Post-hoc tests answer that question while keeping the error rate under control — and Magic Stat reports the statistic each method actually uses, plus a compact letter display that reads like the ones in ecology and agriculture papers.
Running every pairwise t-test after a significant omnibus test inflates the chance of at least one false positive — with several groups, badly. A post-hoc test adjusts the comparisons so that the family-wise error rate is held at the level you chose, at the cost of some power. That trade is the whole point of the method.
The methods differ in what they assume and in the statistic they use. Tukey HSD uses the studentized range with a pooled variance and assumes equal variances and a balanced design. Games-Howell uses the studentized range but does not assume equal variances, so it is the default when the spreads or the sample sizes differ. Bonferroni and Dunn-Sidak are conservative and work in any design. Dunn's test and Conover-Iman are the rank-based methods used after Kruskal-Wallis. Dunnett compares every group only against a control, which has more power than all-pairs when that is the actual question.
Where post-hoc is the wrong choice: when you had specific hypotheses before seeing the data, planned contrasts test them directly and with more power. Discovering the hypothesis after looking at the data and then testing it as if it had been planned is the practice that pre-specification exists to prevent.
Tukey HSD: the studentized range with a pooled variance — equal variances, balanced design.
Games-Howell: the studentized range without the equal-variance assumption — unequal variances or unequal group sizes.
Bonferroni and Dunn-Sidak: conservative corrections for any design.
Dunn's test (with Bonferroni) and Conover-Iman: rank-based, used after Kruskal-Wallis.
Dunnett (vs control): every group compared only against a control — the sharper test when that is the question.
Compact letter display: computed with the Piepho insertion algorithm.
Tukey if the variances are similar and the design is balanced. Games-Howell if the spreads or the group sizes differ, because it does not assume equal variances — the same reasoning as Welch's versus Student's t for two groups.
Because each one carries its own false-positive risk, and with several groups those risks accumulate. A post-hoc method controls the error rate for the whole family of comparisons; that control is the reason the test exists.
Groups that share at least one letter are not significantly different; groups with no letter in common are. It is a compact summary of the pairwise results, and the algorithm used here guarantees that a significant pair never shares a letter — unlike a greedy lettering scheme, which can imply significance the test did not find.
The Means Comparison dialog also carries a batch correction — Benjamini-Hochberg (FDR, the default), Bonferroni, Holm or Hochberg — for the set of tests run together, on top of the within-test post-hoc correction.
No. Everything runs on your machine; the only automatic signal is an anonymous installation counter.