You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何生成DataFrame分组唯一排列对应的子DataFrame集合?

生成指定组合的DataFrame

输入DataFrame

import pandas as pd

input_df = pd.DataFrame({
    'Previous': ['1000', '1000', 'latex', 'latex'], 
    'Ignore':[None, None, ['free'], ['free']], 
    'New': ['100', '200', 'nylon', 'cloth']
})

需求目标

需要生成以下4个DataFrame,每个DataFrame需包含Previous列的所有唯一值('1000'和'latex'),并将New列的不同值做交叉组合:

df1 = pd.DataFrame({
    'Previous': ['1000','latex'],
    'Ignore': [None, ['free']],
    'New': ['100','nylon']
})
df2 = pd.DataFrame({
    'Previous': ['1000','latex'],
    'Ignore': [None, ['free']],
    'New': ['100','cloth']
})
df3 = pd.DataFrame({
    'Previous': ['1000','latex'],
    'Ignore': [None, ['free']],
    'New': ['200','nylon']
})
df4 = pd.DataFrame({
    'Previous': ['1000','latex'],
    'Ignore': [None, ['free']],
    'New': ['200','cloth']
})

解决方案

通过对原DataFrame的行进行组合筛选,仅保留包含所有Previous唯一值的组合,最终得到目标结果:

from itertools import combinations as c

out = [pd.DataFrame(j) for j in c([i[1] for i in input_df.iterrows()], len(input_df['Previous'].unique())) 
       if len(pd.DataFrame(j)['Previous'].unique()) == len(input_df['Previous'].unique())]

代码说明

  • 用itertools.combinations生成原DataFrame行的所有指定长度(等于Previous列唯一值数量)的组合
  • 筛选出组合后Previous列包含所有唯一值的结果,将这些组合转换为DataFrame,最终out列表就是所需的4个目标DataFrame

内容的提问来源于stack exchange,提问作者oliverjenkins

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 10:11:07