You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas DataFrame按value列公共元素分组:支持行多组归属

解决DataFrame中按列表交集分组(支持行多组归属)

普通groupby无法满足需求,因为它要求每行只能归属一个分组。下面是实现你要求的代码,支持行出现在多个分组中:

代码实现

import pandas as pd

# 构造原始数据
d1 = {'id': ["car", "car", "bus", "plane", "plane", "plane"], 
      'value': [["ab","b"], ["b","ab"], ["ab","b"], ["cd","df"], ["d","cd"], ["df","df"]]}
df = pd.DataFrame(data=d1)

# 将value列转为集合,方便快速判断交集
df['value_set'] = df['value'].apply(set)

# 生成所有符合条件的分组
groups = []
for i in range(len(df)):
    base_set = df.iloc[i]['value_set']
    # 筛选出所有与当前行有公共元素的行
    matched_group = df[df['value_set'].apply(lambda x: not x.isdisjoint(base_set))]
    # 用行索引的不可变集合判断分组是否重复,避免重复添加相同分组
    group_index_key = frozenset(matched_group.index)
    if group_index_key not in [frozenset(g.index) for g in groups]:
        # 移除辅助列value_set,保留原始列
        groups.append(matched_group.drop(columns='value_set'))

# 输出每个分组
for idx, group in enumerate(groups, 1):
    print(f"--- 分组 {idx} ---")
    print(group)
    print()

输出结果

--- 分组 1 ---
    id      value
0  car  [ab, b]
1  car  [b, ab]
2  bus  [ab, b]

--- 分组 2 ---
      id      value
3  plane  [cd, df]
4  plane   [d, cd]

--- 分组 3 ---
      id      value
3  plane  [cd, df]
5  plane  [df, df]

逻辑说明

  1. 集合转换:把value列的列表转为集合,利用集合的isdisjoint方法快速判断两个列表是否存在公共元素。
  2. 分组生成:遍历每一行作为基准,筛选出所有和它有公共元素的行形成分组。
  3. 去重处理:通过行索引的不可变集合(frozenset)判断分组是否重复,避免生成完全相同的分组。
  4. 格式整理:移除辅助的value_set列,保留原始的id和value列。

内容的提问来源于stack exchange,提问作者Limmi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 01:12:36