GAM analysis

Generalized additive models without the mgcv boilerplate

A GAM is the answer when the effect of a predictor is not a straight line but you do not want to guess a polynomial. Magic Stat fits a penalized cubic regression spline to the smooth term, chooses the smoothing by GCV/UBRE the way mgcv does, and gives you the curve with its uncertainty.

The honest part

What a GAM is, and when it beats a straight line

A generalized additive model is a GLM in which the effect of a predictor is allowed to curve. Instead of one coefficient per variable, the model fits a smooth s(x) — here a cubic regression spline with a penalty on curvature — and lets the data decide how wiggly that curve may be. The penalty is what keeps it from chasing noise: the smoothing parameter λ is selected by GCV when the scale is estimated (gaussian, Gamma) or UBRE when it is fixed (poisson, binomial), the same criterion mgcv uses with method="GCV.Cp".

The numbers to read are the effective degrees of freedom (edf) of the smooth — near 1 means the curve is essentially a line — and the smooth's own test. Magic Stat reports edf, edf1, ref.df, the Wood (2013) smooth test with the exact Davies p, the deviance explained and the adjusted R², then draws s(x) with ±2SE bands.

It is not always the right model. If the relationship is plausibly linear, a standard GLM is simpler, easier to report and its coefficient is directly interpretable. If your data are clustered or repeated, the dependence belongs in a mixed model or GEE, not in a smoother. And the reproduction caveat is worth stating: current mgcv defaults to REML, which selects a slightly different λ — to match Magic Stat in R you set method="GCV.Cp".

What it does

What Magic Stat gives you for GAM

How it works

From outcome and smooth to the curve

  1. Load the data. one row per observation; the outcome, the smooth variable and any parametric predictors are ordinary columns.
  2. Choose the response and family. the family's canonical link is applied — counts suggest poisson, 0/1 suggests binomial.
  3. Set the smooth term. pick the variable to smooth and the basis size k (leave it at 10 unless you have a reason).
  4. Add parametric terms. optional numeric or categorical predictors that enter with a linear effect.
  5. Read and export. edf and the smooth's p tell you whether the curve is real; the plot goes to the gallery and the model to the report.
Options

The settings, out in the open

Families: gaussian · poisson · binomial · Gamma — canonical link in each case.

Smooth: one s(var) term (cubic regression spline, bs = "cr") · basis size k (3–50, default 10).

Smoothing selection: GCV when the scale is estimated · UBRE when it is fixed — mgcv's method = "GCV.Cp".

Parametric predictors: any mix of numeric and categorical columns; categorical entry uses treatment contrasts with the first category as reference.

Frequently asked

GAM questions, answered honestly

GAM or a linear model?

If the effect is plausibly a straight line, use a linear model or GLM: the coefficient is simpler to report and easier to interpret. Reach for a GAM when the data show curvature you cannot justify with a polynomial, and read the smooth's edf — near 1 means the smoother collapsed towards a line.

Why GCV and not REML?

Magic Stat follows mgcv's method = "GCV.Cp". Modern mgcv defaults to REML, which selects a slightly different λ, so to reproduce a Magic Stat fit in R you set method = "GCV.Cp" explicitly. The results are close, not identical, and the difference is stated rather than hidden.

How do I know the smooth is real?

The smooth test (Wood 2013, exact Davies p) answers whether s(x) explains variation beyond a constant, and edf tells you how curved it is. A significant smooth with edf close to 1 is barely distinguishable from a line.

Can I include categorical predictors?

Yes, alongside the smooth and with treatment contrasts (first category alphabetically as reference), the coding R uses. Non-canonical links are not offered; each family uses its canonical link only.

Does the analysis run locally?

Yes. The fit runs on your computer; nothing but an anonymous installation counter leaves the machine.

Generalized linear models → Nonlinear regression → Mixed models →