You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将两个DataFrame的同类别值随机打乱后填充至第三个DataFrame?

按类别打乱填充DataFrame值的实现方案

问题描述

现有两个结构一致的DataFrame df1 和 df2,索引范围1至n(示例为1-2),数据如下:

df1 = pd.DataFrame({
    'Fruit': ['Apple', 'Pineapple', 'Apple', 'Pineapple'],
    'Indices': [1, 1, 2, 2],
    'Value': [10, 20, 30, 40]
})

df2 = pd.DataFrame({
    'Fruit': ['Apple', 'Pineapple', 'Apple', 'Pineapple'],
    'Indices': [1, 1, 2, 2],
    'Value': [50, 60, 70, 80]
})

另有一个规模为前两者两倍的DataFrame df3,索引范围1至2n,数据如下:

df3 = pd.DataFrame({
    'Fruit': ['Apple', 'Pineapple', 'Apple', 'Pineapple', 'Apple', 'Pineapple', 'Apple', 'Pineapple'],
    'Indices': [1, 1, 2, 2, 3, 3, 4, 4],
    'Value': [np.nan, np.nan, np.nan, np.nan, np.nan, np.nan, np.nan, np.nan]
})

需求:针对每个Fruit类别,将df1和df2中该类别的所有Value值随机打乱后,填充至df3对应Fruit的Value列中(例如Apple类可用值为10、30、50、70,Pineapple类为20、40、60、80,均需打乱填充)。

实现步骤与代码

核心思路

  1. 合并df1和df2的所有数据,按Fruit类别分组
  2. 对每组的Value列进行随机打乱
  3. 通过布尔掩码定位df3中对应类别的行,将打乱后的值批量填充

完整代码

import pandas as pd
import numpy as np

# 初始化示例数据
df1 = pd.DataFrame({
    'Fruit': ['Apple', 'Pineapple', 'Apple', 'Pineapple'],
    'Indices': [1, 1, 2, 2],
    'Value': [10, 20, 30, 40]
})

df2 = pd.DataFrame({
    'Fruit': ['Apple', 'Pineapple', 'Apple', 'Pineapple'],
    'Indices': [1, 1, 2, 2],
    'Value': [50, 60, 70, 80]
})

df3 = pd.DataFrame({
    'Fruit': ['Apple', 'Pineapple', 'Apple', 'Pineapple', 'Apple', 'Pineapple', 'Apple', 'Pineapple'],
    'Indices': [1, 1, 2, 2, 3, 3, 4, 4],
    'Value': [np.nan, np.nan, np.nan, np.nan, np.nan, np.nan, np.nan, np.nan]
})

# 合并数据并按类别打乱Value
combined_df = pd.concat([df1, df2])
shuffled_data = combined_df.groupby('Fruit')['Value'].apply(lambda x: x.sample(frac=1).reset_index(drop=True))

# 填充df3
for fruit in shuffled_data.index.unique():
    # 获取当前类别打乱后的值列表
    target_values = shuffled_data.loc[fruit].tolist()
    # 定位df3中当前类别的行
    fruit_mask = df3['Fruit'] == fruit
    # 批量填充
    df3.loc[fruit_mask, 'Value'] = target_values

print(df3)

代码说明

  • pd.concat([df1, df2]):合并两个DataFrame,得到所有类别的原始值集合
  • groupby('Fruit')['Value'].apply(lambda x: x.sample(frac=1).reset_index(drop=True)):按Fruit分组后,用sample(frac=1)对每组值进行全量随机采样(即打乱顺序),重置索引保证后续填充的顺序正确
  • df3.loc[fruit_mask, 'Value'] = target_values:通过布尔掩码精准定位df3中对应类别的行,批量填充打乱后的值

内容的提问来源于stack exchange,提问作者user6346482

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 08:33:22