R与SPSS对数线性分析参数存在差异的原因咨询
Why Parameter Estimates Differ Between R and SPSS Log-Linear Analyses
Great question! The slight discrepancies in your parameter estimates, plus the coding difference for gender, stem from key defaults in how R and SPSS handle factor coding, reference groups, and model fitting. Let’s break this down clearly:
1. Factor Coding & Reference Group Mismatches
Your R code explicitly sets custom factor levels, while SPSS uses its default coding rules—this is the biggest driver of differences:
- R setup: You defined:
genderlevels asc("female","male")→ treatment coding uses female as the reference group (all gender-related parameters represent differences relative to female)departmentlevels asc("two","one")→ department two is the reference groupadmissionlevels asc("admitted","notadmitted")→ admitted is the reference group
- SPSS setup: You noted SPSS codes
male=0andfemale=1. By default, SPSS uses dummy coding where the lowest-valued category (0, in this case male) is the reference group. This reverses the reference for gender compared to your R model. Similarly, SPSS likely uses the original alphabetical/input order fordepartmentandadmissionas references, which may not match your reversed levels in R.
These flipped reference groups directly change the sign and interpretation of parameter estimates (e.g., a parameter for "female vs male" in SPSS is the inverse of "male vs female" in R).
2. Delta Adjustment in SPSS
Your SPSS syntax includes /CRITERIA=DELTA(.5). This is a default SPSS adjustment that adds a small value (0.5 here) to all cell counts, even if there are no zero frequencies. The goal is to avoid numerical issues with log transformations, but it slightly alters the underlying data used for fitting.
R’s glm(poisson) does not apply this adjustment by default—it uses the raw counts directly. This tiny shift in input data leads to minor differences in parameter estimates.
3. Convergence & Iteration Defaults
- SPSS uses
ITERATE(20)andCONVERGE(0.001)as stopping criteria for model fitting. - R’s
glm()defaults to a stricter convergence tolerance (epsilon=1e-8) and more iterations (25 by default).
When models are near convergence, these differing thresholds can result in slightly different final parameter values, even when the core model structure matches.
How to Align the Results
To make your R and SPSS outputs match closely:
- In R, set reference groups to match SPSS (e.g.,
gender = factor(gender, levels = c("male","female"))to use male as reference) - In SPSS, disable the delta adjustment with
/CRITERIA=DELTA(0) - Ensure both tools use the same convergence criteria
内容的提问来源于stack exchange,提问作者user1205901 - Слава</think_never_used_51bce0c785ca2f68081bfa7d91973934>

