You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Python判断4个历史样本中哪个对总体的代表性最优

Python实现多类型变量样本代表性校验方案

核心逻辑

可以通过自动识别变量类型匹配对应检验方法,批量完成校验后按综合得分选出最优样本,核心判断规则为:检验的p值越大,样本与总体的分布差异越不显著,代表性越好。

依赖准备

需要用到pandas处理数据,scipy.stats做统计检验,安装命令:pip install pandas scipy

完整实现代码

import pandas as pd
from scipy.stats import ks_2samp, chi2_contingency

def evaluate_sample_representativeness(df_population, df_sample, categorical_cols=None, continuous_cols=None):
    """
    评估单样本对总体的代表性
    返回每个变量的检验p值和综合得分
    """
    # 未手动指定变量类型时自动按dtype拆分
    if not categorical_cols:
        categorical_cols = df_population.select_dtypes(include=['object', 'category']).columns.tolist()
    if not continuous_cols:
        continuous_cols = df_population.select_dtypes(include=['float', 'int']).columns.tolist()
    
    result = {}
    # 分类变量用Pearson卡方检验
    for col in categorical_cols:
        # 对齐总体和样本的类别,避免检验报错
        pop_freq = df_population[col].value_counts().sort_index()
        sample_freq = df_sample[col].value_counts().sort_index()
        all_cats = pop_freq.index.union(sample_freq.index)
        pop_freq = pop_freq.reindex(all_cats, fill_value=0)
        sample_freq = sample_freq.reindex(all_cats, fill_value=0)
        contingency_table = pd.DataFrame([pop_freq, sample_freq]).values
        chi2, p, dof, expected = chi2_contingency(contingency_table)
        result[col] = p
    
    # 连续变量用Kolmogorov-Smirnov检验
    for col in continuous_cols:
        pop_data = df_population[col].dropna()
        sample_data = df_sample[col].dropna()
        stat, p = ks_2samp(pop_data, sample_data)
        result[col] = p
    
    # 综合得分默认取所有变量p值的均值,可自定义加权规则
    result['综合得分'] = sum(result.values())/len(result)
    return result

# ------------------- 示例调用 -------------------
# 替换为你的实际数据路径
df_pop = pd.read_csv("你的总体数据路径.csv")
sample_list = [
    pd.read_csv("样本1.csv"),
    pd.read_csv("样本2.csv"),
    pd.read_csv("样本3.csv"),
    pd.read_csv("样本4.csv")
]
sample_names = ["样本1", "样本2", "样本3", "样本4"]

# 批量评估所有样本
all_result = []
for name, sample in zip(sample_names, sample_list):
    res = evaluate_sample_representativeness(df_pop, sample)
    res['样本名称'] = name
    all_result.append(res)

# 按综合得分降序排列,排在第一位的就是代表性最优的样本
result_df = pd.DataFrame(all_result).sort_values('综合得分', ascending=False).reset_index(drop=True)
print(result_df)

注意事项

  • 卡方检验要求列联表中期望频数小于5的格子占比不超过20%,如果分类变量类别多、样本量小,可以把占比低于5%的类别合并为「其他」后再做检验,避免结果偏差
  • 如果有业务优先级更高的变量,可以修改综合得分的计算逻辑,给重要变量的p值设置更高权重,更贴合实际需求
  • 如果所有样本的p值普遍小于0.05,说明所有样本和总体都存在显著分布差异,建议检查抽样逻辑是否存在系统偏差

内容的提问来源于stack exchange,提问作者deepankar srigyan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 23:27:03