PLS-DA is the supervised method for many correlated variables and few samples — the classic situation in metabolomics and chemometrics. Magic Stat runs it from a normal spreadsheet, reports VIP scores and a confusion matrix, and gives you a cross-validated accuracy next to the optimistic fit one.
PLS discriminant analysis fits partial least squares regression against a dummy matrix of the group labels. The components are chosen to maximise the covariance between the measured variables and the groups, which is why it works well when the variables are many, correlated and noisy — the situation where ordinary discriminant analysis becomes unstable.
The trap is that PLS-DA can fit almost any grouping if you give it enough components, and a resubstitution accuracy will look excellent while meaning very little. That is why this dialog reports two accuracies side by side: the fit accuracy and a cross-validated one from a stratified k-fold split. The cross-validated number is the one to believe.
The other useful output is the importance of each variable. The VIP score condenses a variable's contribution across the components; the usual heuristic is that VIP above roughly one marks a variable worth attention. It is a ranking device, not a significance test — there is no p-value attached to a VIP, and this dialog does not pretend there is.
PLS-DA assumes there are groups to separate and that the variables carry enough signal to separate them. If your sample size is small and your variables are many, the honest move is to treat the model as exploratory and to validate it on independent data. When you need orthogonal partial least squares (OPLS-DA), permutation testing of the model, or formal variable selection, R's mixOmics package is the appropriate tool.
Grouping variable: a categorical column with two or more levels. Variables: two or more numeric columns.
Number of components: one to eight, capped by the data at min(n−1, p). The variables are standardised internally by the algorithm, and the cross-validation uses a stratified split with a fixed seed, so the reported accuracy is reproducible.
Discriminant analysis assumes normal predictors within groups and equal covariance, and is stable only when you have many more samples than variables. PLS-DA handles many correlated variables and small samples, which is why it dominates metabolomics and chemometrics. Use discriminant analysis when its assumptions hold.
The fit accuracy is resubstitution — computed on the data the model was trained on — and is optimistic. The cross-validated accuracy comes from a stratified k-fold split and is the more honest estimate of how the model would perform on new samples. Report both, and lead with the second.
As few as gets the job done. More components always fit the training data better and often generalise worse. Increase the number only while the cross-validated accuracy improves, and state how you chose it.
No. VIP is a ranking of variable importance with no p-value. Values around or above one are conventionally treated as worth attention, but they do not establish that a variable matters. Permutation testing of VIP, or formal selection with cross-validation, is beyond this dialog — mixOmics in R is the tool for it.
No. Everything runs on your machine. The only automatic signal is an anonymous installation counter.
Discriminant analysis → PCA analysis → Logistic regression →