关于MASS包lda$scaling与SPSS线性判别结果差异的技术咨询
Hey there, let's break down why you're seeing discrepancies between your R LDA output and the SPSS demo from class—this is a super common gotcha, and it all ties back to how each tool normalizes discriminant function coefficients.
The Core Difference
First, let's unpack what lda$scaling actually gives you. As the docs note, it's the matrix for converting observations to discriminant functions, normalized so the within-group covariance matrix is spherical. Here's how that contrasts with SPSS:
- SPSS outputs raw (unstandardized) discriminant coefficients that directly reflect the original variable scales. These coefficients are optimized to maximize group separation without any extra normalization step.
- MASS's
lda$scalingapplies a normalization: it effectively scales each variable by the inverse square root of its within-group variance before computing the discriminant functions. This makes the within-group covariance matrix spherical (diagonal with 1s), but it changes the coefficient values compared to SPSS.
Data Standardization vs. LDA Coefficient Normalization
You’re totally right that data standardization (like z-scoring) doesn’t change the group separation or the structure of the discriminant functions—LDA is invariant to linear transformations of variables. But here, we’re talking about a different kind of normalization applied directly to the coefficients by R’s lda() function, not pre-processing your data. This is why you’re seeing different numbers even if your data matches the SPSS demo.
How to Match SPSS’s Results in R
If you want to replicate the raw coefficients from SPSS, you can reverse R’s normalization manually. Here’s a step-by-step with code:
- Calculate the within-group covariance matrix from your LDA model.
- Extract the within-group standard deviations (square root of the diagonal of the covariance matrix).
- Multiply each column of
lda$scalingby the corresponding within-group standard deviation to undo the spherical normalization.
library(MASS) # Fit your LDA model (replace with your data and formula) lda_fit <- lda(GroupVariable ~ ., data = your_dataset) # Compute within-group covariance matrix within_cov <- with(lda_fit, withinss / (n - ngroups)) # Get within-group standard deviations within_sd <- sqrt(diag(within_cov)) # Adjust scaling to match SPSS's raw coefficients spss_coefficients <- t(t(lda_fit$scaling) * within_sd)
Bottom Line
The difference isn’t a mistake—it’s just a design choice between the two tools. The underlying discriminant functions (how well they separate groups) are identical; only the coefficient values differ because of the normalization step. Once you adjust the scaling as above, your R results should align with the SPSS demo.
内容的提问来源于stack exchange,提问作者Jay Schyler Raadt

