如何用选中项列表过滤含列表列的Pandas DataFrame
Pandas高效过滤含指定食材的行
需求说明
现有一个Pandas DataFrame,其中Ingredients列的每个元素为对应行的食材列表;另有一个选中食材列表checked_items(示例值:['Carrot', 'Celery', 'Onion'])。需要移除所有Ingredients列列表与checked_items无任何匹配项的行,只要两者存在任意匹配就保留该行。
示例输入:
checked_items=['Carrot', 'Celery', 'Onion'] Col_1 Col_2 Ingredients "a" "e" [Carrot, Ginger, Curry] "b" "f" [Butter, Shallots] "c" "g" [Celery, Onion, Sage, Thyme]
期望输出:
Col_1 Col_2 Ingredients "a" "e" [Carrot, Ginger, Curry] "c" "g" [Celery, Onion, Sage, Thyme]
现有问题代码
你尝试了以下代码,但存在结果不准确、大数据量下效率极低的问题:
mask=[] for ingredient_list in df['Ingredients'].to_list(): if not ingredient_list: mask.append(False) continue i=0 try: for ingredient in ingredient_list: for checked_item in checked_items: if checked_item == ingredient: mask.append(True) raise StopIteration i=i+1 if i==len(categories): mask.append(False) except StopIteration: continue filtered_df = df[mask]
高效解决方案
核心优化思路
- 将
checked_items转为集合,集合的成员查询操作是O(1),远快于列表的O(n) - 使用Pandas的
apply方法结合生成器表达式,实现向量化的条件判断,比手动嵌套Python循环效率高很多 - 自动处理空列表的情况(空列表与集合无交集,直接返回False)
实现代码
# 将选中项转为集合,提升查询效率 checked_set = set(checked_items) # 生成过滤掩码:判断每个食材列表是否与选中集合有交集 mask = df['Ingredients'].apply(lambda lst: any(item in checked_set for item in lst)) # 过滤得到结果 filtered_df = df[mask]
进一步优化(超大数据集)
如果你的DataFrame规模极大(百万行以上),可以用pd.explode结合isin实现更高效的向量化操作:
# 展开Ingredients列,每行一个食材 exploded = df.explode('Ingredients') # 筛选出包含选中食材的行的原始索引 valid_indices = exploded[exploded['Ingredients'].isin(checked_items)].index.unique() # 根据索引过滤原DataFrame filtered_df = df.loc[valid_indices]
这种方法完全利用Pandas的向量化操作,避免了Python层面的循环,在大数据量下性能更优。
内容的提问来源于stack exchange,提问作者Stoggy88
相关产品推荐
相关产品推荐

