You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何gt summary与chisq.test的卡方检验结果存在差异?

卡方检验结果差异原因解析

问题描述

先后执行两次卡方检验得到不同结果:

  1. 使用gt summary实现:
table1_demographic_data <- subsetted.demographic %>% 
  tbl_summary(by= post_covid_dic, 
              missing = "no", missing_text= "NA", 
              type= list(comorbid_HTN ~ "dichotomous"), 
              statistic = list(all_continuous() ~ "{mean}?{sd}", all_categorical() ~ "{n} ({p}%)"), 
              percent = "row") %>% 
  add_p() %>% 
  modify_header(label = "**Variable**", statistic ~ "**Test Statistic**") %>% 
  modify_fmt_fun(statistic ~ style_sigfig) %>% 
  bold_labels() %>% 
  bold_p() %>% 
  modify_caption("table 1: Demographic characteristics")

table1_demographic_data

结果:$\chi^2$=4.3,p=0.039

  1. 使用chisq.test函数实现:
chisq.test(subsetted.demographic$post_covid_dic, subsetted.demographic$comorbid_HTN)

结果:$\chi^2$=3.8076,p=0.05

差异原因分析

  • 连续性修正设置不同:
    R语言中chisq.test()函数默认对2×2列联表启用连续性修正(参数correct=TRUE),这一修正会减小卡方统计量的数值,同时使p值增大。而gt summary的add_p()函数在处理2×2表的卡方检验时,默认不启用连续性修正(等同于调用chisq.test(correct=FALSE)),这是两者结果差异的核心原因。
    验证这一点很简单:执行chisq.test(subsetted.demographic$post_covid_dic, subsetted.demographic$comorbid_HTN, correct=FALSE),得到的统计量和p值会与gt summary的输出完全匹配。

  • 结果精度展示差异:
    gt summary输出的卡方值经过了四舍五入处理(显示为4.3),而chisq.test输出的是原始计算的精确值(3.8076),这进一步放大了两者结果的视觉差异,但并非核心原因。

内容的提问来源于stack exchange,提问作者user3315563

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 12:54:57