You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中实现类R的Pairwise Wilcoxon检验并生成对比表格?

在Python中执行Pairwise Wilcoxon检验并生成美观对比表格

我懂你想要的——在Python里复刻R中pairwise.wilcox.test()的功能,还得输出清爽好看的两两对比表格对吧?确实Python没有直接的内置函数一步到位,但用几个常用库组合起来完全能搞定,我给你一步步拆解:

先明确两种核心场景

R的pairwise.wilcox.test()默认处理独立样本(本质是两两Mann-Whitney U检验,和Wilcoxon秩和检验等价);如果是配对样本(比如同一组对象在不同条件下的测量值),则要用Wilcoxon配对符号秩检验。我们分开来实现:

第一步:安装并导入所需库

先确保你有这些工具:

import numpy as np
import pandas as pd
from scipy.stats import mannwhitneyu, wilcoxon
from itertools import combinations
from tabulate import tabulate
from statsmodels.stats.multitest import multipletests
  • scipy:执行统计检验
  • itertools:生成所有组的两两组合
  • tabulate:快速生成美观的文本表格
  • statsmodels:做p值校正(和R的默认行为对齐)

场景1:独立样本的两两检验(对应R默认)

假设你有几组独立的样本数据,我们模拟一组示例数据,然后执行检验:

1. 准备示例数据

np.random.seed(42)  # 固定随机种子保证结果可复现
group_a = np.random.normal(loc=5, scale=1, size=30)
group_b = np.random.normal(loc=6, scale=1, size=30)
group_c = np.random.normal(loc=5.5, scale=1, size=30)
group_d = np.random.normal(loc=7, scale=1, size=30)

# 整理成DataFrame方便处理
data = pd.DataFrame({
    'Group': ['A']*30 + ['B']*30 + ['C']*30 + ['D']*30,
    'Value': np.concatenate([group_a, group_b, group_c, group_d])
})

2. 执行两两Mann-Whitney U检验(等价于Wilcoxon秩和)

# 获取所有组名
groups = data['Group'].unique()
# 生成所有两两组合
pairwise_pairs = list(combinations(groups, 2))

# 存储检验结果
results = []
for g1, g2 in pairwise_pairs:
    # 提取两组数据
    d1 = data[data['Group'] == g1]['Value']
    d2 = data[data['Group'] == g2]['Value']
    # 执行检验(双侧检验)
    stat, p_val = mannwhitneyu(d1, d2, alternative='two-sided')
    # 保存结果
    results.append({
        'Group 1': g1,
        'Group 2': g2,
        'U Statistic': round(stat, 2),
        'p-value': round(p_val, 4)
    })

# 转成DataFrame
results_df = pd.DataFrame(results)

3. 加入p值校正(和R对齐)

R的pairwise.wilcox.test()默认会做Bonferroni校正,我们手动加上:

# 对p值进行Bonferroni校正
corrected_pvals = multipletests(results_df['p-value'], method='bonferroni')[1]
results_df['Corrected p-value'] = np.round(corrected_pvals, 4)
results_df['Significant'] = ['Yes' if p < 0.05 else 'No' for p in corrected_pvals]

场景2:配对样本的两两Wilcoxon检验

如果你的数据是配对的(比如同一批样本在4种不同处理下的测量值),我们用Wilcoxon配对符号秩检验:

1. 准备配对示例数据

np.random.seed(42)
paired_data = pd.DataFrame({
    'Sample_ID': range(30),
    'Condition_A': np.random.normal(loc=5, scale=1, size=30),
    'Condition_B': np.random.normal(loc=5.5, scale=1, size=30),
    'Condition_C': np.random.normal(loc=6, scale=1, size=30),
    'Condition_D': np.random.normal(loc=6.5, scale=1, size=30)
})

2. 执行两两配对检验

# 获取所有条件组名
condition_groups = paired_data.columns[1:]  # 排除Sample_ID列
paired_pairs = list(combinations(condition_groups, 2))

paired_results = []
for c1, c2 in paired_pairs:
    d1 = paired_data[c1]
    d2 = paired_data[c2]
    # 执行Wilcoxon配对检验
    stat, p_val = wilcoxon(d1, d2, alternative='two-sided')
    paired_results.append({
        'Condition 1': c1,
        'Condition 2': c2,
        'Z Statistic': round(stat, 2),
        'p-value': round(p_val, 4)
    })

paired_results_df = pd.DataFrame(paired_results)
# 同样加入p值校正
corrected_pvals_paired = multipletests(paired_results_df['p-value'], method='bonferroni')[1]
paired_results_df['Corrected p-value'] = np.round(corrected_pvals_paired, 4)
paired_results_df['Significant'] = ['Yes' if p < 0.05 else 'No' for p in corrected_pvals_paired]

生成美观的对比表格

1. 文本格式表格(适合终端/日志输出)

用tabulate生成网格状清晰表格:

# 输出独立样本结果
print("=== 独立样本两两Wilcoxon检验结果 ===")
print(tabulate(results_df, headers='keys', tablefmt='grid', showindex=False))

# 输出配对样本结果
print("\n=== 配对样本两两Wilcoxon检验结果 ===")
print(tabulate(paired_results_df, headers='keys', tablefmt='grid', showindex=False))

输出效果是规整的网格表格,一眼就能看清各组的p值和显著性。

2. 带样式的HTML表格(适合报告/网页)

用pandas的style模块生成带高亮的美观表格,还能保存为HTML:

# 独立样本表格样式
styled_independent = results_df.style \
    .highlight_min(subset=['Corrected p-value'], color='#82e0aa')  # 高亮显著的p值
    .format({'p-value': '{:.4f}', 'Corrected p-value': '{:.4f}', 'U Statistic': '{:.2f}'}) \
    .set_caption('独立样本两两Wilcoxon秩和检验结果(Bonferroni校正)')

# 保存为HTML文件
styled_independent.to_html('independent_wilcoxon_results.html')

# 配对样本表格同理
styled_paired = paired_results_df.style \
    .highlight_min(subset=['Corrected p-value'], color='#82e0aa') \
    .format({'p-value': '{:.4f}', 'Corrected p-value': '{:.4f}', 'Z Statistic': '{:.2f}'}) \
    .set_caption('配对样本两两Wilcoxon符号秩检验结果(Bonferroni校正)')

styled_paired.to_html('paired_wilcoxon_results.html')

这样生成的HTML表格会高亮显著的p值,格式清爽,直接嵌入报告或者浏览器打开都很美观。

内容的提问来源于stack exchange,提问作者tea_time

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 23:24:06