如何对比Python列表与Pandas列表列的完全/部分/不匹配行
Pandas非扁平化列表与目标列表的匹配判断方案
实现代码
import pandas as pd # 定义目标列表和待处理DataFrame target_list = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10] df = pd.DataFrame({ "col1": ["a", "b", "c", "d"], "col2": [[1, 2, 3, 4, 5], [20, 30, 100], [1, 3, 20], [20, 30]] }) # 转换为集合提升元素存在性判断效率 target_set = set(target_list) # 初始化结果容器 full_match = [] partial_match = [] no_match = [] # 逐行处理判断匹配类型 for _, row in df.iterrows(): col2_elements = set(row["col2"]) common = col2_elements & target_set if len(common) == len(col2_elements): full_match.append(row["col1"]) elif len(common) > 0: partial_match.append(row["col1"]) else: no_match.append(row["col1"]) # 按指定格式输出结果 print(f'"100% matches are in rows: {", ".join(full_match)}"') print(f'"Partial matches are in rows: {", ".join(partial_match)}"') print(f'"Dismatches are in rows: {", ".join(no_match)}"')
代码说明
- 利用集合的交集运算快速判断元素重叠情况,比逐个遍历列表效率更高
- 匹配规则:
- 完全匹配:
col2中所有元素都存在于目标列表 - 部分匹配:
col2中存在至少一个目标列表的元素,但不全在 - 不匹配:
col2中没有任何元素属于目标列表
- 完全匹配:
- 最后通过
str.join()将结果列表转为符合要求的字符串格式
内容的提问来源于stack exchange,提问作者Mr.Slow
相关产品推荐
相关产品推荐

