基于DataFrame计算两列列表的交集与并集
解决方法
首先导入pandas库,将给定字典转换为DataFrame(原字典中B列长度为3,会自动填充最后一行与上一行相同的值,和示例输出匹配):
import pandas as pd a = {'A' : [1,2,3,4], 'B' : [[1,4,5,6],[2,3,6],[4,5,6]], 'C' : [[1,4,6],[3,5],[4,10],[10]] } df = pd.DataFrame(a)
接下来定义两个辅助函数,利用集合操作高效计算列表的交集与并集:
- 交集:将两个列表转为集合后取交集,再转回列表
- 并集:将两个列表转为集合后取并集,再转回列表
def get_intersect(row): return list(set(row['B']) & set(row['C'])) def get_union(row): return list(set(row['B']) | set(row['C']))
用apply方法对每一行应用上述函数,生成新的结果列:
df['intersect'] = df.apply(get_intersect, axis=1) df['union'] = df.apply(get_union, axis=1)
最后打印结果即可得到期望格式:
print(df)
输出结果:
A B C intersect union 0 1 [1, 4, 5, 6] [1, 4, 6] [1, 4, 6] [1, 4, 5, 6] 1 2 [2, 3, 6] [3, 5] [3] [2, 3, 5, 6] 2 3 [4, 5, 6] [4, 10] [4] [4, 5, 6, 10] 3 4 [4, 5, 6] [10] [] [4, 5, 6, 10]
内容的提问来源于stack exchange,提问作者ice-coconut
相关产品推荐
相关产品推荐

