GEE analysis

Generalized estimating equations without the geepack setup

GEE estimates population-average effects when observations come in clusters — repeated visits, siblings, matched sites — without assuming a full distribution for the random effects. Magic Stat runs the geepack-equivalent fit and reports the robust standard errors that GEE is used for.

The honest part

What GEE estimates, and when a mixed model fits better

GEE solves a set of estimating equations that depend on a mean model (a GLM with the family and link you choose) and a working correlation structure describing how observations inside a cluster move together. Its output is the population-average effect: how the average response changes with a predictor, across the whole population. Inference uses a sandwich (robust) estimator, so it stays valid even if the working correlation is not exactly right.

This is a different question from a mixed model's. A random-intercept mixed model gives a subject-specific effect (what happens to a given subject); GEE gives the marginal effect (what happens on average). For a linear model they coincide; for logistic and other non-linear links they do not, and choosing the wrong one is a real mistake, not a nuance.

The engine mirrors geepack's geeglm: the same Gauss–Seidel β → γ → α iteration, the same moment estimators for the correlation, the same convergence tolerance. It also reports when the fit did not converge — which the plain R summary does not do for you.

What it does

What Magic Stat gives you for GEE

How it works

From clusters to robust standard errors

  1. Load the data. one row per observation, with a column identifying the cluster (subject, site, matched set).
  2. Choose the outcome and the cluster column. the dialog reports how many clusters there are and their sizes before the fit.
  3. Set family, link and working correlation. independence is the safe default; exchangeable suits repeated measures without a natural order.
  4. Choose predictors. numeric or categorical.
  5. Run and read. coefficients with robust SEs, the correlation parameters, and a note if the fit did not converge.
Options

The settings, out in the open

Families and links: gaussian · poisson · binomial · Gamma, with the links listed per family above.

Working correlation: independence · exchangeable · AR(1) · unstructured.

Clusters: defined by an ID column; the data are ordered by it before fitting, as geepack requires.

Scale: estimated by moments for gaussian and Gamma, fixed at the GLM value for poisson and binomial — the geepack convention.

Frequently asked

GEE questions, answered honestly

GEE or a mixed model?

GEE answers the population-average question and stays valid under a wrong working correlation when clusters are numerous. A mixed model answers the subject-specific question and can be preferable with few clusters or when you want variance components. For non-linear links the two effects genuinely differ; the choice is about which question you are asking.

Which working correlation should I use?

independence is the conservative default and, with a robust sandwich, still gives valid inference. exchangeable fits repeated measures with no time order; AR(1) assumes evenly spaced, ordered measurements; unstructured is the most flexible and the most parameter-hungry.

What is the difference from an ordinary GLM?

A GLM assumes independent rows. If your rows are clustered, its standard errors are too small. GEE accounts for the within-cluster dependence through the working correlation and the robust variance.

Does the app warn about non-convergence?

Yes — unlike the plain R summary, Magic Stat shows a note when the iteration did not converge, so a non-converged fit is not presented as a clean result.

Is anything sent to a server?

No. The estimation runs locally; only an anonymous installation counter is sent.

Mixed models → Generalized linear models → Logistic regression →