You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于分组与概率从另一DataFrame随机添加行的实现方法

问题描述

现有三个Pandas DataFrame:

  1. Main:包含Region、Type、Value列,存在大量重复的Region/Type组合行,示例:
Region  Type    Value
A        1       600
A        2       700
A        2       750
B        1       700
B        1       500
B        2       900
B        2      1000
  1. Prob:包含Region、Type、Probability列,作为Main的概率查询表,为每个Region/Type组合指定固定概率,示例:
Region  Type   Probability
A        1       20%
A        2       30%
B        1       40%
B        2       90%
  1. Extras:结构与Main一致,行数更多,示例:
Region  Type   Value
A        1       600
A        1       300
A        2       700
A        2       950
B        1       700
B        1        50
B        2       900
B        2      1200
B        2       300

核心需求

遍历Main的每一行,根据当前行的Region和Type从Prob中匹配对应概率,随机判断是否要从Extras中抽取同Region/Type的一行添加到Main;抽取后需将该行从Extras中移除,避免重复抽取。实际场景中Main约10000行,Extras约40000行,概率约20%,预期新增约2000行。

实现方案

步骤1:预处理数据

先把Prob的百分比概率转为小数,同时将Extras按Region/Type分组,用字典存储每组的行索引列表,方便快速随机抽取并移除。

步骤2:遍历执行逻辑

对Main的每一行匹配对应概率,生成随机数判断是否触发添加操作;如果触发,从Extras对应分组中随机选一行,添加到Main后移除该行。

完整代码示例

import pandas as pd
import numpy as np

# 替换为你的真实数据
main_df = pd.DataFrame({
    'Region': ['A', 'A', 'A', 'B', 'B', 'B', 'B'],
    'Type': [1, 2, 2, 1, 1, 2, 2],
    'Value': [600, 700, 750, 700, 500, 900, 1000]
})

prob_df = pd.DataFrame({
    'Region': ['A', 'A', 'B', 'B'],
    'Type': [1, 2, 1, 2],
    'Probability': ['20%', '30%', '40%', '90%']
})

extras_df = pd.DataFrame({
    'Region': ['A', 'A', 'A', 'A', 'B', 'B', 'B', 'B', 'B'],
    'Type': [1, 1, 2, 2, 1, 1, 2, 2, 2],
    'Value': [600, 300, 700, 950, 700, 50, 900, 1200, 300]
})

# 预处理:将百分比转为小数
prob_df['Probability'] = prob_df['Probability'].str.replace('%', '').astype(float) / 100

# 预处理Extras:按Region+Type分组存储索引列表
extras_groups = {}
for (region, type_val), group in extras_df.groupby(['Region', 'Type']):
    extras_groups[(region, type_val)] = group.index.tolist()

# 批量收集要添加的行
rows_to_add = []

# 遍历Main每一行执行逻辑
for _, row in main_df.iterrows():
    key = (row['Region'], row['Type'])
    # 获取对应概率
    prob = prob_df[(prob_df['Region'] == key[0]) & (prob_df['Type'] == key[1])]['Probability'].iloc[0]
    
    # 随机判断是否添加
    if np.random.rand() < prob:
        # 检查分组是否还有可用行
        if key in extras_groups and len(extras_groups[key]) > 0:
            # 随机选一行索引
            selected_idx = np.random.choice(extras_groups[key])
            # 取出该行并加入待添加列表
            selected_row = extras_df.loc[selected_idx].copy()
            rows_to_add.append(selected_row)
            # 从分组中移除已选索引
            extras_groups[key].remove(selected_idx)

# 批量合并新增行到Main
main_df = pd.concat([main_df, pd.DataFrame(rows_to_add)], ignore_index=True)

# 更新Extras:保留未被抽取的行
remaining_indices = []
for idx_list in extras_groups.values():
    remaining_indices.extend(idx_list)
extras_df = extras_df.loc[remaining_indices].reset_index(drop=True)

性能优化说明

  • 用字典存储Extras分组索引,避免每次抽取时重复过滤DataFrame,适配万级数据量的高效查询。
  • 批量收集新增行后一次性合并,比逐行追加更节省内存和时间。
  • 最后一次性过滤Extras剩余行,减少多次修改DataFrame的开销。

内容的提问来源于stack exchange,提问作者Richard Dixon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 23:15:14