You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对比DataFrame的字符串列表列,提取Col B中不在Col A的元素生成Col C?

高效生成Pandas列表列的差异列

首先是你提供的DataFrame:

import pandas as pd

d = {'Col A': [['Singapore','Germany','UK'],['Ireland','Japan','Australia'],['India','Korea','Vietnam']], 
     'Col B': [['Singapore','Germany','UK'],['Ireland','Japan'],['India','Mexico','Argentina']]}

df = pd.DataFrame(data=d)

需求明确:生成新列Col C,仅包含存在于Col B但不在Col A的元素,其余情况返回空列表。

两种简便实现方式:

1. 保留Col B元素顺序的写法

如果需要严格保持Col B中元素的原始顺序,直接用列表推导式遍历Col B元素,判断是否不在Col A中:

df['Col C'] = df.apply(lambda row: [item for item in row['Col B'] if item not in row['Col A']], axis=1)

2. 更简洁的集合差集写法(不保证顺序)

利用集合的差集特性,一行代码搞定,缺点是会打乱元素原有顺序:

df['Col C'] = df.apply(lambda row: list(set(row['Col B']) - set(row['Col A'])), axis=1)

验证结果

执行后df的内容如下:

Col ACol BCol C
['Singapore', 'Germany', 'UK']['Singapore', 'Germany', 'UK'][]
['Ireland', 'Japan', 'Australia']['Ireland', 'Japan'][]
['India', 'Korea', 'Vietnam']['India', 'Mexico', 'Argentina']['Mexico', 'Argentina']

(注:集合差集写法的Col C元素顺序可能因集合特性有所不同)

内容的提问来源于stack exchange,提问作者Michael Kessler

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 03:35:17