请求编写Chi-Squared Test函数:计算p值并检验原假设
自定义卡方检验(Chi-Squared Test)函数实现
核心逻辑说明
卡方检验用于判断两组分类特征是否存在关联,原假设为两组特征相互独立,备择假设为两组特征存在显著关联。下面的函数手动实现核心计算逻辑,每一步附带注释,方便理解:
import math from scipy.stats import chi2 # 若无法使用scipy,可替换为自定义的卡方CDF计算(见备注) def chi_squared_test(observed): """ 卡方检验函数,计算统计量、p值并检验原假设 参数: observed (list of lists): 二维列表,代表列联表的观测频数 示例: [[男-购买数, 男-未购买数], [女-购买数, 女-未购买数]] 返回: tuple: (chi2_statistic, p_value, conclusion) conclusion为字符串,说明是否拒绝原假设 """ # 1. 计算行和、列和、总频数 row_sums = [sum(row) for row in observed] col_sums = [sum(col) for col in zip(*observed)] total = sum(row_sums) # 2. 计算期望频数矩阵 expected = [] for i in range(len(row_sums)): row_expected = [] for j in range(len(col_sums)): exp = (row_sums[i] * col_sums[j]) / total row_expected.append(exp) expected.append(row_expected) # 3. 计算卡方统计量 chi2_stat = 0.0 for i in range(len(observed)): for j in range(len(observed[i])): if expected[i][j] == 0: continue # 避免除以0,实际中需确保期望频数≥5,否则检验不可靠 chi2_stat += ((observed[i][j] - expected[i][j]) ** 2) / expected[i][j] # 4. 计算自由度 df = (len(row_sums) - 1) * (len(col_sums) - 1) # 5. 计算p值:卡方分布的生存函数(1 - CDF) p_value = chi2.sf(chi2_stat, df) # 6. 假设检验(默认显著性水平α=0.05) alpha = 0.05 if p_value < alpha: conclusion = f"p值({p_value:.4f}) < α({alpha}),拒绝原假设:两组特征存在显著关联" else: conclusion = f"p值({p_value:.4f}) ≥ α({alpha}),无法拒绝原假设:两组特征相互独立" return chi2_stat, p_value, conclusion # ---------------------- 示例使用 ---------------------- if __name__ == "__main__": # 示例列联表:行=性别(男、女),列=是否购买(是、否) observed_data = [ [25, 15], # 男:25人购买,15人未购买 [18, 22] # 女:18人购买,22人未购买 ] stat, p, res = chi_squared_test(observed_data) print(f"卡方统计量: {stat:.4f}") print(f"p值: {p:.4f}") print(res)
备注:无scipy时的p值计算替代方案
如果无法使用scipy.stats.chi2,可以用卡方分布的正态近似(自由度df≥30时适用),或者手动实现伽马函数计算CDF:
def gamma_function(x): # 伽马函数近似(斯特林公式) if x < 0.5: return math.pi / (math.sin(math.pi * x) * gamma_function(1 - x)) x -= 1 ln_gamma = x * math.log(x + 4.5) - (x + 4.5) + 0.5 * math.log(2 * math.pi) ln_gamma += (1/(12*(x+1)) - 1/(360*(x+1)**3) + 1/(1260*(x+1)**5)) return math.exp(ln_gamma) def chi2_cdf(x, df): # 卡方分布CDF计算(正则化伽马函数) if x <= 0: return 0.0 k = df / 2.0 t = x / 2.0 incomplete_gamma = 0.0 term = math.exp(-t) * (t ** (k - 1)) incomplete_gamma += term for i in range(1, 100): term *= t / (k + i - 1) incomplete_gamma += term incomplete_gamma /= gamma_function(k) return incomplete_gamma # 此时p值计算替换为: # p_value = 1 - chi2_cdf(chi2_stat, df)
使用注意事项
- 列联表中每个单元格的期望频数需≥5,否则检验结果不可靠,此时建议合并类别或使用Fisher精确检验
- 显著性水平α可根据需求修改(默认0.05)
内容的提问来源于stack exchange,提问作者Sadegh
相关产品推荐
相关产品推荐

