You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R与SPSS对数线性分析参数存在差异的原因咨询

Why Parameter Estimates Differ Between R and SPSS Log-Linear Analyses

Great question! The slight discrepancies in your parameter estimates, plus the coding difference for gender, stem from key defaults in how R and SPSS handle factor coding, reference groups, and model fitting. Let’s break this down clearly:


1. Factor Coding & Reference Group Mismatches

Your R code explicitly sets custom factor levels, while SPSS uses its default coding rules—this is the biggest driver of differences:

  • R setup: You defined:
    • gender levels as c("female","male") → treatment coding uses female as the reference group (all gender-related parameters represent differences relative to female)
    • department levels as c("two","one") → department two is the reference group
    • admission levels as c("admitted","notadmitted") → admitted is the reference group
  • SPSS setup: You noted SPSS codes male=0 and female=1. By default, SPSS uses dummy coding where the lowest-valued category (0, in this case male) is the reference group. This reverses the reference for gender compared to your R model. Similarly, SPSS likely uses the original alphabetical/input order for department and admission as references, which may not match your reversed levels in R.

These flipped reference groups directly change the sign and interpretation of parameter estimates (e.g., a parameter for "female vs male" in SPSS is the inverse of "male vs female" in R).

2. Delta Adjustment in SPSS

Your SPSS syntax includes /CRITERIA=DELTA(.5). This is a default SPSS adjustment that adds a small value (0.5 here) to all cell counts, even if there are no zero frequencies. The goal is to avoid numerical issues with log transformations, but it slightly alters the underlying data used for fitting.

R’s glm(poisson) does not apply this adjustment by default—it uses the raw counts directly. This tiny shift in input data leads to minor differences in parameter estimates.

3. Convergence & Iteration Defaults

  • SPSS uses ITERATE(20) and CONVERGE(0.001) as stopping criteria for model fitting.
  • R’s glm() defaults to a stricter convergence tolerance (epsilon=1e-8) and more iterations (25 by default).

When models are near convergence, these differing thresholds can result in slightly different final parameter values, even when the core model structure matches.


How to Align the Results

To make your R and SPSS outputs match closely:

  • In R, set reference groups to match SPSS (e.g., gender = factor(gender, levels = c("male","female")) to use male as reference)
  • In SPSS, disable the delta adjustment with /CRITERIA=DELTA(0)
  • Ensure both tools use the same convergence criteria

内容的提问来源于stack exchange,提问作者user1205901 - Слава</think_never_used_51bce0c785ca2f68081bfa7d91973934>

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 11:17:38