为何R与Python的Logistic模型置信区间结果存在差异?
问题:Logistic回归模型在R与Python中置信区间结果差异的原因
我正在构建Logistic回归模型,用于预测个体健康状况(0=良好,1=不佳)是否受二手烟暴露及其他人口统计学因素的影响。
R中的模型实现
我在R中使用以下代码构建模型(变量已重命名以标识类型):
shs_health_model <- glm(health_binary ~ continuous1 + continuous2 + continuous3 + binary1 + binary2 + binary3 + binary_smoke, data=mydata, family="binomial")
提取优势比(OR)、置信区间和p值的代码:
shs_coef <- exp(cbind(OR = coef(shs_health_model), confint(shs_health_model))) cbind(shs_coef, P = summary(shs_health_model$coefficients[,'Pr(>|z|)']))
R的输出结果:
| 变量 | OR | 2.5 % | 97.5 % | P |
|---|---|---|---|---|
| (Intercept) | 4.3165924 | 0.40965543 | 56.1120993 | 0.239687729 |
| continuous1 | 1.0351381 | 0.99352565 | 1.0799027 | 0.101217153 |
| continuous2 | 0.5532695 | 0.31446430 | 0.9486117 | 0.032469282 |
| continuous3 | 0.9718181 | 0.92737903 | 1.0195214 | 0.231590917 |
| binary1 | 6.8387667 | 1.23306396 | 41.0086896 | 0.027803143 |
| binary2 | 0.1796848 | 0.05899218 | 0.4867787 | 0.001274915 |
| binary3 | 0.6038520 | 0.16843900 | 1.8854196 | 0.405517065 |
| binary_smoke | 0.5601202 | 0.14253453 | 2.4350002 | 0.412941761 |
Python中的模型实现
合作者在Python中使用statsmodels.api的Logit函数,基于相同变量和数据集构建模型:
model = sm.Logit(y, X).fit()
提取优势比、置信区间和p值的代码:
import numpy as np import pandas as pd import statsmodels.api as sm odds_ratios = np.exp(model.params) conf = np.exp(model.conf_int()) p_values = model.pvalues logistic_results['total_environ_index_cat'] = pd.DataFrame({ 'Odds Ratios': odds_ratios, 'Confidence Intervals (Lower)': conf[0], 'Confidence Intervals (Upper)': conf[1], 'p-values': p_values })
Python的输出结果(可见部分变量的置信区间值存在差异,但OR和p值与R的输出一致):
| OR | CI(low) | CI(high) | p | |
|---|---|---|---|---|
| const | 4.317 | 0.377 | 49.42 | 0.24 |
| continuous1 | 1.035 | 0.993 | 1.079 | 0.101 |
| continuous2 | 0.553 | 0.322 | 0.952 | 0.032 |
| continuous3 | 0.972 | 0.927 | 1.018 | 0.232 |
| binary1 | 6.839 | 1.233 | 37.918 | 0.028 |
| binary2 | 0.18 | 0.063 | 0.511 | 0.001 |
| binary3 | 0.604 | 0.184 | 1.982 | 0.406 |
| binary_smoke | 0.56 | 0.14 | 2.243 | 0.413 |
疑问
为何两者的输出会存在差异?并非其中一方的置信区间始终更宽,且据我所知两者计算的均为95%置信区间。
内容的提问来源于stack exchange,提问作者AIRE Study
相关产品推荐
相关产品推荐

