如何基于指定列筛选pandas DataFrame中不在给定列表内的行
实现方法
你可以通过以下两种常见方式实现需求:
方法一:逐行打包元组匹配(小数据量场景推荐,写法简单)
直接把col1和col2两列的值打包成元组,再用isin方法取反筛选即可,完整代码如下:
import pandas as pd data = [ {'col1': 11, 'col2': 111, 'col3': 1111}, {'col1': 22, 'col2': 222, 'col3': 2222}, {'col1': 33, 'col2': 333, 'col3': 3333}, {'col1': 44, 'col2': 444, 'col3': 4444} ] lst = [(11, 111), (22, 222), (99, 999)] df = pd.DataFrame(data) # 核心筛选逻辑:打包col1、col2为元组,判断不在lst中 res_df = df[~df[['col1', 'col2']].apply(tuple, axis=1).isin(lst)] # 转换为要求的字典列表格式 output = res_df.to_dict('records') print(output)
方法二:关联过滤(大数据量场景推荐,性能更高)
如果你已经把过滤列表转成了DataFrame,可以用左连接+标记位的方式实现反匹配,避免逐行处理的性能损耗:
list_df = pd.DataFrame(lst, columns=['col1', 'col2']) # 左连接并添加匹配标记 merged = df.merge(list_df, on=['col1', 'col2'], how='left', indicator=True) # 仅保留只在原df中存在的行,删除标记列 res_df = merged[merged['_merge'] == 'left_only'].drop(columns='_merge') output = res_df.to_dict('records')
两种方法输出的结果都与你给出的预期一致:
[{'col1': 33, 'col2': 333, 'col3': 3333}, {'col1': 44, 'col2': 444, 'col3': 4444}]
内容的提问来源于stack exchange,提问作者dina
相关产品推荐
相关产品推荐

