You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于集合交集筛选DataFrame:保留df1中与df2有交集的行

解决方法

步骤1:构造示例DataFrame

先还原你提供的两个DataFrame:

import pandas as pd

df1 = pd.DataFrame({
    'categories': [['hello','world'], ['gogo','albert'], ['dodo']]
})

df2 = pd.DataFrame({
    'categories': [['hello','world'], ['albert'], ['dodji']]
})

步骤2:实现筛选逻辑

核心是检查df1每一行的categories列表,是否与df2中任意一行的categories列表存在交集。为提升效率,先把df2的所有列表转换成集合:

# 将df2的categories转换为集合列表,减少交集计算开销
df2_category_sets = [set(lst) for lst in df2['categories'].tolist()]

# 定义检查函数:判断当前行列表是否与df2中任意集合有交集
def check_intersection(row_cats):
    row_set = set(row_cats)
    for cat_set in df2_category_sets:
        if row_set & cat_set:
            return True
    return False

# 对df1执行筛选
df_all = df1[df1['categories'].apply(check_intersection)]

步骤3:查看结果

执行后df_all的输出为:

categories
0  [hello, world]
1  [gogo, albert]

完全符合预期的筛选结果。

内容的提问来源于stack exchange,提问作者michel gold

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 20:30:32