You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用少量代码对Pandas多列与Product执行独立双样本t检验?

独立双样本t检验:Product与多列的对比实现

你可以通过循环遍历目标列的方式,用少量代码完成所有列的独立双样本t检验,避免重复写冗余代码。以下是完整实现:

数据集与依赖导入

import pandas as pd
from scipy.stats import ttest_ind

# 构建数据集
data = {
    'Product': ['laptop', 'printer','printer','printer','laptop','printer','laptop','laptop','printer','printer'],
    'Purchase_cost': [120.09, 150.45, 300.12, 450.11, 200.55,175.89,124.12,113.12,143.33,375.65],
    'Warranty_years':[3,2,2,1,4,1,2,3,1,2],
    'service_cost': [5,5,10,4,7,10,4,6,12,3]
}

df = pd.DataFrame(data)

完整检验代码

# 按Product分组
laptop_group = df[df['Product'] == 'laptop']
printer_group = df[df['Product'] == 'printer']

# 定义需要检验的列
target_columns = ['Purchase_cost', 'Warranty_years', 'service_cost']

# 遍历列执行t检验并输出结果
for col in target_columns:
    t_stat, p_value = ttest_ind(laptop_group[col], printer_group[col])
    print(f"=== {col} 与 Product 的独立双样本t检验 ===")
    print(f"t统计量: {t_stat:.4f}")
    print(f"p值: {p_value:.4f}\n")

代码说明

  • 先完成分组:将数据按Product分为laptop和printer两组,避免重复分组操作
  • 用列表存储需要检验的列名,通过循环批量处理,减少冗余代码
  • 每次检验后输出清晰的结果,包括t统计量和p值,方便解读

内容的提问来源于stack exchange,提问作者nasa313

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 03:37:52