When several independent tests address the same question, combining their p-values into one answer is a small, well-defined calculation — and one that is easy to do wrong. Magic Stat offers the two standard methods, Fisher's and Stouffer's, with the same conventions as the R metap package.
Fisher's method turns the k p-values into a chi-square statistic: χ² = −2·Σ ln(pᵢ), with 2k degrees of freedom. It reacts strongly to a single very small p-value, so it is sensitive to one strong study. Stouffer's method converts each pᵢ to a normal deviate zᵢ = Φ⁻¹(1−pᵢ) and combines them, z = Σzᵢ/√k; it can be weighted, z = Σwᵢzᵢ/√(Σwᵢ²), when studies deserve different importance.
Both methods assume the tests being combined are independent. That is the foundation. If the same data or the same subjects appear in more than one test, the combined p-value is not valid, and neither method is the right answer — no combination of dependent p-values is offered here.
One convention is worth knowing: the Stouffer combination here is one-sided (upper tail), matching metap's sumz. A two-sided variant is not offered, and the app reports the one-sided result explicitly rather than leaving you to guess.
Methods: Fisher (χ², df = 2k) · Stouffer (Z, unweighted or weighted).
Input: p-values in (0, 1], one per row; p = 1 is allowed and behaves as in R.
Weights: positive numbers, the same length as the p-values, Stouffer only.
Sidedness: one-sided (upper tail), the metap convention.
Fisher's method is governed by the smallest p-value and is the classic choice for a handful of studies. Stouffer's method is more balanced and lets you weight studies by precision or importance. Neither is universally better; the decision should match how you think about the individual tests.
No. Both methods assume the p-values come from independent tests. Combining correlated p-values (overlapping samples, repeated measures, the same data reanalysed) invalidates the result. If your tests are dependent, these methods are not the answer, and the app offers no shortcut around that.
The reference implementation, metap's sumz, is one-sided, and Magic Stat follows it. Presenting a two-sided p under the same name would silently disagree with the package you might cite.
Yes — Fisher's method can produce a very small combined p from several moderately small ones, and that is a feature of the method, not an error. It is also why a single extreme p-value can dominate.
No. The arithmetic runs locally; the only automatic signal is an anonymous installation counter.