You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对含多列不平衡唯一值的DataFrame带替换采样,平衡分布且保原长度?

解决DataFrame多列类别平衡采样(保持原行数)

针对你需要将包含A、B、C、D、E列的DataFrame进行带替换采样,使B/C/D/E列每个唯一值行数相同且采样后总长度与原数据一致的需求,以下是具体实现方案:

核心思路

通过计算每列每个类别的目标采样次数,采用加权带替换采样的方式,优先选择能填补各列类别缺失次数的行,最终生成满足多列平衡且行数匹配的数据集。

步骤1:计算基础参数

先获取原数据的总行数,以及每列的唯一值分组:

import pandas as pd
import numpy as np

# 假设df是你的原始DataFrame
n_total = len(df)
target_cols = ['B', 'C', 'D', 'E']

# 存储每列的唯一值对应数据子集
col_groups = {}
for col in target_cols:
    col_groups[col] = {val: group for val, group in df.groupby(col)}
    col_groups[col]['n_unique'] = len(col_groups[col]) - 1  # 记录该列唯一值数量

步骤2:确定每列类别目标采样次数

按总行数均分每列的唯一值,余数随机分配给部分类别,保证总次数等于原数据行数:

col_target_counts = {}
for col in target_cols:
    n_unique = col_groups[col]['n_unique']
    base_count = n_total // n_unique
    remainder = n_total % n_unique
    
    # 生成每个唯一值的目标采样次数
    target_counts = {}
    for i, val in enumerate(col_groups[col].keys() - {'n_unique'}):
        target_counts[val] = base_count + 1 if i < remainder else base_count
    col_target_counts[col] = target_counts

步骤3:加权带替换采样

通过迭代采样,每次选择能最大程度填补各列类别缺失次数的行,直到达到目标行数:

sampled_rows = []
remaining_counts = {col: counts.copy() for col, counts in col_target_counts.items()}

while len(sampled_rows) < n_total:
    # 计算每行的采样权重:与对应列类别的剩余需求次数成正比
    weights = []
    for _, row in df.iterrows():
        weight = 1
        for col in target_cols:
            val = row[col]
            weight *= remaining_counts[col][val]
        weights.append(weight)
    
    # 归一化权重,避免数值过大
    weights = np.array(weights)
    if weights.sum() == 0:
        # 若剩余需求为0,随机采样剩余行数
        remaining = n_total - len(sampled_rows)
        sampled_rows.extend(df.sample(n=remaining, replace=True).to_dict('records'))
        break
    weights = weights / weights.sum()
    
    # 带替换采样一行
    sampled_idx = np.random.choice(df.index, p=weights)
    sampled_row = df.loc[sampled_idx]
    sampled_rows.append(sampled_row)
    
    # 更新剩余需求次数
    for col in target_cols:
        val = sampled_row[col]
        remaining_counts[col][val] -= 1
        if remaining_counts[col][val] < 0:
            remaining_counts[col][val] = 0

# 转换为最终的平衡DataFrame
balanced_df = pd.DataFrame(sampled_rows).reset_index(drop=True)

验证采样结果

检查每列的唯一值行数分布是否符合要求:

for col in target_cols:
    print(f"{col}列行数分布:")
    print(balanced_df[col].value_counts())

内容的提问来源于stack exchange,提问作者Kaihua Hou

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 20:35:15