You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas DataFrame中tuple列与指定列表匹配生成布尔列的最优方法

错误原因说明

  • 第一种写法df['found'] = any((True for x in test_list if x in df['colors_new']))的问题:any()返回的是单个布尔值,会将整个列赋值为同一个结果,无法实现逐行判断。
  • 第二种写法df['found'] = ['red' in x for x in df['colors_new']]的问题:仅硬编码了单个匹配规则,无法动态适配test_list中的多个待匹配元素。

最优实现方案

前置优化

首先将test_list转为集合,成员查询的时间复杂度从列表的O(n)降至O(1),匹配元素越多性能提升越明显:

test_set = set(test_list)

1. 通用场景最高效写法(列表推导式,适合绝大多数中小数据集)

列表推导式是pandas处理行级可迭代对象性能最高的方案,比apply快30%~200%:

df['found'] = [any(color in test_set for color in row_tuple) for row_tuple in df['colors_new']]

也可以用集合交集判断,写法更简洁:

df['found'] = [bool(test_set & set(row_tuple)) for row_tuple in df['colors_new']]

2. 超大规模数据集写法(完全向量化操作)

如果数据量超过10万行,推荐用explode+分组判断的向量化方案,避免Python层循环:

df['found'] = df['colors_new'].explode().isin(test_set).groupby(level=0).any()

样例输出

执行后得到的df结果如下:

numberscolorscolors_newfound
1[red, blue](red, blue)True
2[white](white,)False
3[blue, yellow](blue, yellow)True

内容的提问来源于stack exchange,提问作者Manza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 15:06:06