为何相同数据下stats.chisquare与stats.chi2_contingency结果不同?
问题:scipy中chisquare与chi2_contingency结果不一致的原因
我有一个存储9个数字组计数的数组data_count,用本福特定律生成了预期计数数组expected_counts,想检验二者分布是否一致。但用scipy的stats.chisquare和stats.chi2_contingency函数后得到的结果差异显著,而Scipy文档表明二者在合适场景下应返回相同结果。以下是代码及输出:
import numpy as np from scipy import stats data_count = [34, 10, 8, 16, 14, 5, 4, 7, 4] expected_counts = [31, 18, 13, 10, 8, 7, 6, 5, 5] expected_percentage=[(i/sum(expected_counts))*100 for i in expected_counts] data_percentage=[(i/sum(data_count))*100 for i in data_count] # method 1 res1 = stats.chisquare(f_obs=data_percentage, f_exp=expected_percentage) print(res1.pvalue) # method 2 combined = np.array([data_count, expected_counts]) res2 = stats.chi2_contingency(combined, correction=False) print(res2.pvalue)
输出结果:
0.04329908403353834 0.45237501133745583
问题原因及修正
1. chisquare参数使用错误
stats.chisquare是拟合优度检验工具,要求f_obs和f_exp传入原始观测计数,而非百分比。你将计数转换为百分比后,相当于对数据做了等比例缩放,计算出的卡方统计量完全偏离正确值,导致p值错误。
2. chi2_contingency使用场景错误
chi2_contingency是用来做列联表独立性检验的,它的输入应该是实际观测的列联表(比如不同分组下各类型的频数),而不是把观测计数和预期计数当成两行放入数组。这个函数的核心是检验行列变量是否独立,而非观测与预期分布的拟合程度,你的用法完全不符合该函数的设计目的。
修正后的代码
import numpy as np from scipy import stats data_count = [34, 10, 8, 16, 14, 5, 4, 7, 4] expected_counts = [31, 18, 13, 10, 8, 7, 6, 5, 5] # 正确使用拟合优度检验:传入原始计数 res = stats.chisquare(f_obs=data_count, f_exp=expected_counts, correction=False) print("正确的拟合优度检验p值:", res.pvalue)
内容的提问来源于stack exchange,提问作者Shuxiang Cai
相关产品推荐
相关产品推荐

