You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

卡方拟合优度检验对所有乘数均判定拟合成功,问题出在哪?

问题描述

我在用卡方拟合优度检验,将观测数据数组与含自由参数(乘数形式)的理论数据数组匹配,以找到能实现最佳拟合的乘数取值。但检验结果显示所有乘数均拟合成功,这显然不合理,因此编写了测试脚本排查问题:

import numpy as np
import scipy.stats as stats

multiplier = [-250,-10,-3,1,4,8,23,78,950,1000000]
a = [1,1,1,1,1,1,1]
b = [4,4,4,4,4,4,4]
n = len(multiplier)
chi_p_values = np.zeros(n)
myp = 0.01
foundfit = 0

def adapt_for_chi(a,b): # a = array of observed frequencies, b = array of theoretical expected frequencies for the proposed distribution 
    return [(np.sum(b)/np.sum(a))*i for i in a] # This way, sum(a) and sum(b) will be the same (requirement for stats.chisquare)

for i in range(n):
    a_multiplied = [multiplier[i]*freq for freq in a]
    a_multiplied_for_chi = adapt_for_chi(a_multiplied,b)
    chi_square_test_statistic,p_value = stats.chisquare(a_multiplied_for_chi,b) 
    chi_p_values[i] = p_value
    if p_value > myp:
        foundfit += 1

print('p-values: ',chi_p_values)
print('Successful tests: ',foundfit)

预期仅当乘数为4时拟合成功,对应输出Successful tests: 1,但实际输出为:

p values:  [1. 1. 1. 1. 1. 1. 1. 1. 1. 1.]
Successful tests:  10

请问操作中存在什么错误?


问题分析与解决

你的核心错误出在adapt_for_chi函数的逻辑上,它直接抵消了乘数的所有差异,同时卡方检验的使用逻辑也有误:

  1. adapt_for_chi函数抹掉了乘数的影响
    不管你给a乘什么数值,经过这个函数处理后,a_multiplied_for_chi的每个元素都会被计算为:
    (sum(b)/sum(a_multiplied)) * (multiplier[i] * 1)
    其中sum(a_multiplied)是7 * multiplier[i],sum(b)是28,代入后每个元素结果都是(28/(7*multiplier[i])) * multiplier[i] = 4,和理论数组b的每个元素完全一致。这就导致无论用哪个乘数,处理后的观测数组和理论数组完全相同,卡方统计量为0,p值自然是1.0,所有测试都被判定为“成功”。

  2. 卡方检验的前置逻辑错误
    你不需要强行将观测数组缩放到和理论数组总和一致——这个操作完全违背了用乘数调整观测数据来匹配理论数据的初衷。正确的做法是直接用乘数调整后的观测数组和理论数组做检验,同时注意卡方检验的前提:所有观测和理论频数必须非负,且理论频数建议大于5。

修正后的代码示例:

import numpy as np
import scipy.stats as stats

multiplier = [-250,-10,-3,1,4,8,23,78,950,1000000]
a = [1,1,1,1,1,1,1]
b = [4,4,4,4,4,4,4]
n = len(multiplier)
chi_p_values = np.zeros(n)
myp = 0.01
foundfit = 0

for i in range(n):
    a_multiplied = np.array([multiplier[i] * freq for freq in a])
    # 跳过负乘数的情况,频数不能为负
    if (a_multiplied < 0).any():
        chi_p_values[i] = 0.0
        continue
    # 直接用调整后的观测数组和理论数组做卡方检验
    chi_square_test_statistic, p_value = stats.chisquare(a_multiplied, b)
    chi_p_values[i] = p_value
    if p_value > myp:
        foundfit += 1

print('p-values: ', chi_p_values)
print('Successful tests: ', foundfit)

修正后,只有当乘数为4时,a_multiplied与b完全一致,卡方统计量为0,p值1.0;其他乘数下,a_multiplied和b差异显著,p值远小于0.01,最终foundfit会等于1,符合预期。

内容的提问来源于stack exchange,提问作者Wild Feather

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 14:37:30