分组执行Chi-squared Test的报错与结果差异问题咨询
卡方检验分组执行问题的解决方案
问题背景
现有交叉表数据集,需按indicator分组执行卡方检验(避免循环):
table2 <- data.frame( indicator = c("A_3_respondent_hohh", "A_3_respondent_hohh", "B_2_hh_hosting_displaced_persons", "B_2_hh_hosting_displaced_persons", "o7_current_settlement_type", "o7_current_settlement_type", "o7_current_settlement_type" ), var = c("no", "yes", "no", "yes", "city_smt", "regional_center", "village"), n_female = c(3L, 3L, 5L, 1L, 3L, 1L, 2L), n_male = c(NA, 9L, 6L, 3L, 3L, 6L, NA), n_Overall = c(3L, 12L, 11L, 4L, 6L, 7L, 2L) )
使用以下代码处理时,针对A_3_respondent_hohh指标报错:
tbl2 <- table2 |> select(indicator, var, starts_with("n_"), -c(n_Overall)) |> mutate(across(starts_with("n_"), ~replace_na(.x, 0))) |> group_by(indicator) |> mutate(var = as.factor(var)) |> summarise(chi_st = chisq.test(n_male, n_female)$statistic, chi_p = chisq.test(n_male, n_female)$p.value)
错误信息:
Error in `summarise()`: ! Problem while computing `chi_st = chisq.test(n_male, n_female)$statistic`. ℹ The error occurred in group 1: indicator = "A_3_respondent_hohh". Caused by error in `chisq.test()`: ! 'x' and 'y' must have at least 2 levels
但单独对该指标执行检验可正常运行:
t2 <- table2 |> filter(indicator == "A_3_respondent_hohh") |> select(indicator, var, starts_with("n_"), -c(n_Overall)) |> mutate(across(starts_with("n_"), ~replace_na(.x, 0))) chisq.test(t2[, c("n_male", "n_female")])
需解决两个问题:
- 分组执行时的报错问题
- 分组与单独执行时统计量存在细微差异的原因
注:已知Fisher精确检验更适合,但需使用卡方检验。
报错原因与解决方法
报错根源
分组代码中chisq.test(n_male, n_female)的用法错误:该调用将两个向量视为分类变量的观测值,而非列联表的频数。对于A_3_respondent_hohh指标,替换NA后的n_female向量为c(3,3),仅包含1个水平,不符合chisq.test对输入变量至少2个水平的要求,因此触发报错。
而单独执行时,传入的是频数矩阵(t2[, c("n_male", "n_female")]),这是卡方检验处理列联表的正确用法,因此可正常运行。
修正后的代码
在分组汇总时,将n_male和n_female组合成频数矩阵传入chisq.test,同时仅调用一次检验函数以避免随机差异:
library(dplyr) tbl2 <- table2 |> select(indicator, var, starts_with("n_"), -n_Overall) |> mutate(across(starts_with("n_"), ~replace_na(.x, 0))) |> group_by(indicator) |> summarise( # 构造频数矩阵并执行卡方检验,将结果存入列表 chi_test = list(chisq.test(cbind(n_male, n_female))), # 提取统计量和p值 chi_statistic = chi_test[[1]]$statistic, chi_p_value = chi_test[[1]]$p.value ) |> select(-chi_test) # 移除临时存储的检验结果
统计量细微差异的原因
- 调用逻辑不同:原分组代码将向量当作观测值传入
chisq.test,而单独执行时传入的是频数矩阵,两种场景下卡方检验的计算逻辑完全不同,导致统计量出现差异。 - 重复调用的随机波动:原代码中两次调用
chisq.test,若检验过程涉及Monte Carlo模拟(如期望频数较小时),每次调用的随机种子不同,会导致统计量和p值产生细微波动。修正后的代码仅调用一次检验函数,可消除此类差异。
内容的提问来源于stack exchange,提问作者Hanna Kurovska
相关产品推荐
相关产品推荐

