Combine p-values

Combine independent p-values with Fisher and Stouffer

When several independent tests address the same question, combining their p-values into one answer is a small, well-defined calculation — and one that is easy to do wrong. Magic Stat offers the two standard methods, Fisher's and Stouffer's, with the same conventions as the R metap package.

The honest part

Two methods, one assumption

Fisher's method turns the k p-values into a chi-square statistic: χ² = −2·Σ ln(pᵢ), with 2k degrees of freedom. It reacts strongly to a single very small p-value, so it is sensitive to one strong study. Stouffer's method converts each pᵢ to a normal deviate zᵢ = Φ⁻¹(1−pᵢ) and combines them, z = Σzᵢ/√k; it can be weighted, z = Σwᵢzᵢ/√(Σwᵢ²), when studies deserve different importance.

Both methods assume the tests being combined are independent. That is the foundation. If the same data or the same subjects appear in more than one test, the combined p-value is not valid, and neither method is the right answer — no combination of dependent p-values is offered here.

One convention is worth knowing: the Stouffer combination here is one-sided (upper tail), matching metap's sumz. A two-sided variant is not offered, and the app reports the one-sided result explicitly rather than leaving you to guess.

What it does

What Magic Stat gives you for combining p-values

How it works

From a column of p-values to one answer

  1. Choose the method. Fisher or Stouffer.
  2. Enter the p-values. one per row; anywhere from one to many.
  3. Add weights if needed. only for Stouffer, and only positive values.
  4. Run. the combined statistic (χ² or z), its df where relevant, and the combined p.
  5. Read the verdict. the result is compared with α = 0.05 and written into the report with the method stated.
Options

The settings, out in the open

Methods: Fisher (χ², df = 2k) · Stouffer (Z, unweighted or weighted).

Input: p-values in (0, 1], one per row; p = 1 is allowed and behaves as in R.

Weights: positive numbers, the same length as the p-values, Stouffer only.

Sidedness: one-sided (upper tail), the metap convention.

Frequently asked

Combining p-values, answered honestly

Fisher or Stouffer?

Fisher's method is governed by the smallest p-value and is the classic choice for a handful of studies. Stouffer's method is more balanced and lets you weight studies by precision or importance. Neither is universally better; the decision should match how you think about the individual tests.

Can the tests be dependent?

No. Both methods assume the p-values come from independent tests. Combining correlated p-values (overlapping samples, repeated measures, the same data reanalysed) invalidates the result. If your tests are dependent, these methods are not the answer, and the app offers no shortcut around that.

Why is there no two-sided Stouffer option?

The reference implementation, metap's sumz, is one-sided, and Magic Stat follows it. Presenting a two-sided p under the same name would silently disagree with the package you might cite.

Can a combined p be smaller than every input p?

Yes — Fisher's method can produce a very small combined p from several moderately small ones, and that is a feature of the method, not an error. It is also why a single extreme p-value can dominate.

Is my data uploaded?

No. The arithmetic runs locally; the only automatic signal is an anonymous installation counter.

Meta-analysis → Power analysis → Medical statistics →