You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

请求编写Chi-Squared Test函数:计算p值并检验原假设

自定义卡方检验(Chi-Squared Test)函数实现

核心逻辑说明

卡方检验用于判断两组分类特征是否存在关联,原假设为两组特征相互独立,备择假设为两组特征存在显著关联。下面的函数手动实现核心计算逻辑,每一步附带注释,方便理解:

import math
from scipy.stats import chi2  # 若无法使用scipy,可替换为自定义的卡方CDF计算(见备注)

def chi_squared_test(observed):
    """
    卡方检验函数,计算统计量、p值并检验原假设
    
    参数:
        observed (list of lists): 二维列表,代表列联表的观测频数
                                  示例: [[男-购买数, 男-未购买数], [女-购买数, 女-未购买数]]
    
    返回:
        tuple: (chi2_statistic, p_value, conclusion)
               conclusion为字符串,说明是否拒绝原假设
    """
    # 1. 计算行和、列和、总频数
    row_sums = [sum(row) for row in observed]
    col_sums = [sum(col) for col in zip(*observed)]
    total = sum(row_sums)
    
    # 2. 计算期望频数矩阵
    expected = []
    for i in range(len(row_sums)):
        row_expected = []
        for j in range(len(col_sums)):
            exp = (row_sums[i] * col_sums[j]) / total
            row_expected.append(exp)
        expected.append(row_expected)
    
    # 3. 计算卡方统计量
    chi2_stat = 0.0
    for i in range(len(observed)):
        for j in range(len(observed[i])):
            if expected[i][j] == 0:
                continue  # 避免除以0,实际中需确保期望频数≥5,否则检验不可靠
            chi2_stat += ((observed[i][j] - expected[i][j]) ** 2) / expected[i][j]
    
    # 4. 计算自由度
    df = (len(row_sums) - 1) * (len(col_sums) - 1)
    
    # 5. 计算p值:卡方分布的生存函数(1 - CDF)
    p_value = chi2.sf(chi2_stat, df)
    
    # 6. 假设检验(默认显著性水平α=0.05)
    alpha = 0.05
    if p_value < alpha:
        conclusion = f"p值({p_value:.4f}) < α({alpha}),拒绝原假设:两组特征存在显著关联"
    else:
        conclusion = f"p值({p_value:.4f}) ≥ α({alpha}),无法拒绝原假设:两组特征相互独立"
    
    return chi2_stat, p_value, conclusion

# ---------------------- 示例使用 ----------------------
if __name__ == "__main__":
    # 示例列联表:行=性别(男、女),列=是否购买(是、否)
    observed_data = [
        [25, 15],  # 男:25人购买,15人未购买
        [18, 22]   # 女:18人购买,22人未购买
    ]
    stat, p, res = chi_squared_test(observed_data)
    print(f"卡方统计量: {stat:.4f}")
    print(f"p值: {p:.4f}")
    print(res)

备注:无scipy时的p值计算替代方案

如果无法使用scipy.stats.chi2,可以用卡方分布的正态近似(自由度df≥30时适用),或者手动实现伽马函数计算CDF:

def gamma_function(x):
    # 伽马函数近似(斯特林公式)
    if x < 0.5:
        return math.pi / (math.sin(math.pi * x) * gamma_function(1 - x))
    x -= 1
    ln_gamma = x * math.log(x + 4.5) - (x + 4.5) + 0.5 * math.log(2 * math.pi)
    ln_gamma += (1/(12*(x+1)) - 1/(360*(x+1)**3) + 1/(1260*(x+1)**5))
    return math.exp(ln_gamma)

def chi2_cdf(x, df):
    # 卡方分布CDF计算(正则化伽马函数)
    if x <= 0:
        return 0.0
    k = df / 2.0
    t = x / 2.0
    incomplete_gamma = 0.0
    term = math.exp(-t) * (t ** (k - 1))
    incomplete_gamma += term
    for i in range(1, 100):
        term *= t / (k + i - 1)
        incomplete_gamma += term
    incomplete_gamma /= gamma_function(k)
    return incomplete_gamma

# 此时p值计算替换为:
# p_value = 1 - chi2_cdf(chi2_stat, df)

使用注意事项

  • 列联表中每个单元格的期望频数需≥5,否则检验结果不可靠,此时建议合并类别或使用Fisher精确检验
  • 显著性水平α可根据需求修改(默认0.05)

内容的提问来源于stack exchange,提问作者Sadegh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 06:05:26