You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求推荐Python中对比DataFrame两分区各变量的工具包

实用工具包推荐:对比DataFrame分区变量差异

针对你需要按part字段(0/1)批量对比数值、字符变量分区差异的需求,推荐几个现成工具包,不用从零写代码:

1. ydata-profiling(原Pandas Profiling)

  • 核心能力:自动生成交互式数据报告,支持按指定字段分组对比。能自动识别数值/字符变量,批量计算数值变量的KS统计量、相关性,字符变量的类别关联指标(含Cramer's V相关逻辑)。
  • 快速上手:指定分组字段为part,报告会直接展示两个分区下每个变量的差异统计,包括分布对比、统计量数值,不用手动遍历计算。
    from ydata_profiling import ProfileReport
    profile = ProfileReport(df, comparisons={"groupby": "part"})
    profile.to_file("partition_comparison.html")
    

2. Feature-engine

  • 核心能力:专注特征工程的工具包,内置的特征选择模块可以批量计算变量与分组字段的关联度,完美匹配你的需求:数值变量用KS统计量,字符变量用Cramer's V。
  • 快速上手:通过SelectBySingleClassPerformance类,指定评分指标后批量输出结果,还能直接筛选出差异显著的变量。
    from feature_engine.selection import SelectBySingleClassPerformance
    
    # 数值变量用KS统计量
    selector_num = SelectBySingleClassPerformance(
        variables=[数值列列表],
        scoring="ks",
        threshold=0.1,  # 可自定义阈值筛选变量
        target="part"
    )
    selector_num.fit(df)
    num_ks_scores = selector_num.performance_
    
    # 字符变量用Cramer's V
    selector_cat = SelectBySingleClassPerformance(
        variables=[字符列列表],
        scoring="cramers_v",
        threshold=0.1,
        target="part"
    )
    selector_cat.fit(df)
    cat_cramers_scores = selector_cat.performance_
    

3. Sweetviz

  • 核心能力:轻量型数据可视化对比工具,生成的报告交互性强,能直观展示两个分区的变量差异:数值变量的分布重叠度(含KS值)、字符变量的类别占比差异,还会自动标记差异显著的变量。
  • 快速上手:直接传入两个分区的数据集,一键生成对比报告。
    import sweetviz as sv
    # 拆分两个分区
    part0 = df[df["part"] == 0]
    part1 = df[df["part"] == 1]
    # 生成对比报告
    comparison_report = sv.compare([part0, "Partition 0"], [part1, "Partition 1"])
    comparison_report.show_html("partition_diff.html")
    

4. SciPy + Pandas 轻量组合

  • 核心能力:如果需要完全自定义统计量计算逻辑(比如要同时输出Pearson相关、Kappa系数),用SciPy提供的原生统计函数配合Pandas批量处理最灵活,代码量小且可控。
  • 快速上手示例:
    import pandas as pd
    from scipy.stats import ks_2samp, chi2_contingency
    from sklearn.metrics import cohen_kappa_score
    
    # 数值变量:计算KS统计量和Pearson相关
    num_cols = df.select_dtypes(include=["int64", "float64"]).columns.drop("part")
    num_results = []
    for col in num_cols:
        ks_stat, ks_pval = ks_2samp(df[df["part"]==0][col], df[df["part"]==1][col])
        pearson_corr = df[[col, "part"]].corr().iloc[0,1]
        num_results.append({"variable": col, "ks_stat": ks_stat, "pearson_corr": pearson_corr})
    num_results_df = pd.DataFrame(num_results)
    
    # 字符变量:计算Cramer's V和Kappa系数
    cat_cols = df.select_dtypes(include=["object", "category"]).columns
    cat_results = []
    for col in cat_cols:
        # Cramer's V
        contingency = pd.crosstab(df[col], df["part"])
        chi2, pval, dof, expected = chi2_contingency(contingency)
        cramers_v = (chi2 / (df.shape[0] * (min(contingency.shape)-1))) ** 0.5
        # Kappa系数(需确保变量是分类编码)
        kappa = cohen_kappa_score(df[col], df["part"])
        cat_results.append({"variable": col, "cramers_v": cramers_v, "kappa": kappa})
    cat_results_df = pd.DataFrame(cat_results)
    

内容的提问来源于stack exchange,提问作者Logic_Problem_42

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 10:40:17