You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas基于列将数据均分至分箱,使各区间Good数量相等?

解决方法:基于Good=1的样本均分区间

要实现每个分数区间的Good数量相等,核心思路是先针对Good=1的样本按数量均分区间,再将该区间规则应用到整个数据集——直接用pd.cut或pd.qcut无法满足需求,因为这两个方法分别按数值范围、全量数据数量分箱,而非针对Good=1的样本数量均分。

步骤1:准备数据

导入库并创建示例DataFrame:

import pandas as pd

data = {
    'Score': [100,100,100,300,400,400,600,600,600,650,650,650,700,770,770,800,890,890],
    'Good': [0,0,0,0,0,0,1,1,0,0,0,1,1,1,1,0,1,1]
}
df = pd.DataFrame(data)

步骤2:基于Good=1的样本划分区间

筛选出所有Good=1的行,按Score排序后,按目标数量均分分组:

# 筛选并排序Good=1的样本
good_samples = df[df['Good'] == 1].sort_values('Score')
# 计算分组数:总Good数 ÷ 每组目标数量(示例中总共有8个1,每组2个,分4组)
total_good = len(good_samples)
target_per_bin = 2
num_bins = total_good // target_per_bin

# 给Good=1的样本按数量均分分组
good_samples['bin_group'] = pd.qcut(good_samples['Score'], q=num_bins, labels=False)

# 提取每组的Score边界,生成区间标签
bin_bounds = good_samples.groupby('bin_group')['Score'].agg(['min', 'max']).reset_index()
bin_labels = []
for idx in range(num_bins):
    if idx < num_bins - 1:
        bin_labels.append(f"{bin_bounds['min'][idx]} - {bin_bounds['max'][idx]}")
    else:
        bin_labels.append(f"> {bin_bounds['min'][idx]}")

步骤3:将区间规则应用到全量数据

用得到的边界对整个DataFrame的Score分箱,再统计每个区间的Good总数:

# 构建分箱边界(包含最小Score和最大Score+1,确保覆盖所有数据)
score_edges = [df['Score'].min()] + list(bin_bounds['max'][:-1]) + [df['Score'].max() + 1]
# 给全量数据分箱
df['bin'] = pd.cut(df['Score'], bins=score_edges, labels=bin_labels, include_lowest=True)

# 统计每个区间的Good数量
result = df.groupby('bin')['Good'].sum().reset_index().rename(columns={'Good': 'Goods'})
print(result)

输出结果

bin  Goods
0  100 - 600      2
1  650 - 700      2
2  770 - 800      2
3       > 890      2

内容的提问来源于stack exchange,提问作者Ash

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 16:46:12