You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

分组执行Chi-squared Test的报错与结果差异问题咨询

卡方检验分组执行问题的解决方案

问题背景

现有交叉表数据集,需按indicator分组执行卡方检验(避免循环):

table2 <-  data.frame(
  indicator = c("A_3_respondent_hohh", "A_3_respondent_hohh", 
                   "B_2_hh_hosting_displaced_persons", "B_2_hh_hosting_displaced_persons", 
                   "o7_current_settlement_type", "o7_current_settlement_type", "o7_current_settlement_type"
), 
var = c("no", "yes", "no", "yes", "city_smt", "regional_center", 
           "village"), 
n_female = c(3L,  3L, 5L, 1L, 3L, 1L, 2L), 
n_male = c(NA, 9L, 6L, 3L, 3L, 6L, NA), 
n_Overall = c(3L, 12L, 11L, 4L, 6L, 7L, 2L)
)

使用以下代码处理时,针对A_3_respondent_hohh指标报错:

tbl2 <- table2 |>
  select(indicator, var, starts_with("n_"), -c(n_Overall)) |>
  mutate(across(starts_with("n_"), ~replace_na(.x, 0))) |>
  group_by(indicator) |>
  mutate(var  = as.factor(var)) |>
  summarise(chi_st = chisq.test(n_male, n_female)$statistic,
            chi_p = chisq.test(n_male, n_female)$p.value)

错误信息:

Error in `summarise()`:
! Problem while computing `chi_st = chisq.test(n_male, n_female)$statistic`.
ℹ The error occurred in group 1: indicator = "A_3_respondent_hohh".
Caused by error in `chisq.test()`:
! 'x' and 'y' must have at least 2 levels

但单独对该指标执行检验可正常运行:

t2 <- table2 |> filter(indicator == "A_3_respondent_hohh") |>
  select(indicator, var, starts_with("n_"), -c(n_Overall)) |>
  mutate(across(starts_with("n_"), ~replace_na(.x, 0)))
chisq.test(t2[, c("n_male", "n_female")])

需解决两个问题:

  1. 分组执行时的报错问题
  2. 分组与单独执行时统计量存在细微差异的原因

注:已知Fisher精确检验更适合,但需使用卡方检验。

报错原因与解决方法

报错根源

分组代码中chisq.test(n_male, n_female)的用法错误:该调用将两个向量视为分类变量的观测值,而非列联表的频数。对于A_3_respondent_hohh指标,替换NA后的n_female向量为c(3,3),仅包含1个水平,不符合chisq.test对输入变量至少2个水平的要求,因此触发报错。

而单独执行时,传入的是频数矩阵(t2[, c("n_male", "n_female")]),这是卡方检验处理列联表的正确用法,因此可正常运行。

修正后的代码

在分组汇总时,将n_male和n_female组合成频数矩阵传入chisq.test,同时仅调用一次检验函数以避免随机差异:

library(dplyr)

tbl2 <- table2 |>
  select(indicator, var, starts_with("n_"), -n_Overall) |>
  mutate(across(starts_with("n_"), ~replace_na(.x, 0))) |>
  group_by(indicator) |>
  summarise(
    # 构造频数矩阵并执行卡方检验,将结果存入列表
    chi_test = list(chisq.test(cbind(n_male, n_female))),
    # 提取统计量和p值
    chi_statistic = chi_test[[1]]$statistic,
    chi_p_value = chi_test[[1]]$p.value
  ) |>
  select(-chi_test) # 移除临时存储的检验结果

统计量细微差异的原因

  1. 调用逻辑不同:原分组代码将向量当作观测值传入chisq.test,而单独执行时传入的是频数矩阵,两种场景下卡方检验的计算逻辑完全不同,导致统计量出现差异。
  2. 重复调用的随机波动:原代码中两次调用chisq.test,若检验过程涉及Monte Carlo模拟(如期望频数较小时),每次调用的随机种子不同,会导致统计量和p值产生细微波动。修正后的代码仅调用一次检验函数,可消除此类差异。

内容的提问来源于stack exchange,提问作者Hanna Kurovska

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 11:50:34