You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

百万行Pandas单列高效匹配水果列表(部分匹配)方案问询

高效实现百万行DataFrame的多子串部分匹配

核心需求

处理百万行单列DataFrame,判断每行内容是否包含fruit_list中的任意子串,同时收集所有匹配的子串,要求性能优于启用regex=True的str.contains方法。


最优方案:Aho-Corasick多模式匹配(pyahocorasick库)

Aho-Corasick算法专门用于多子串批量匹配,预编译所有子串后,每个字符串只需遍历一次,时间复杂度接近线性,适合百万级数据场景。

步骤与代码

  1. 安装依赖库
pip install pyahocorasick
  1. 实现匹配逻辑
import pandas as pd
import ahocorasick

# 示例数据
fruit_list = ["app", "appl", "banana", "pear"]
df = pd.DataFrame({
    "only_col": ["apple", "banana", "cherry"]
})

# 构建AC自动机并加载所有子串
auto = ahocorasick.Automaton()
for s in fruit_list:
    auto.add_word(s, s)
auto.make_automaton()

# 定义匹配函数:返回是否匹配、匹配的子串列表
def find_matches(s):
    matches = set()
    for _, match in auto.iter(s):
        matches.add(match)
    if matches:
        return True, ", ".join(sorted(matches))
    return False, pd.NA

# 应用到DataFrame,生成目标列
df[["inlist", "str_found"]] = df["only_col"].apply(
    lambda x: pd.Series(find_matches(x))
)
  1. 输出结果
only_col  inlist str_found
0    apple    True  app, appl
1   banana    True     banana
2   cherry   False       <NA>

纯Pandas替代方案(无需第三方库)

如果不想安装外部库,可预编译优化后的正则表达式,性能略逊于AC自动机,但优于原生循环:

import pandas as pd
import re

fruit_list = ["app", "appl", "banana", "pear"]
df = pd.DataFrame({
    "only_col": ["apple", "banana", "cherry"]
})

# 预编译正则(转义特殊字符避免匹配异常)
pattern = re.compile("|".join(re.escape(s) for s in fruit_list))

def find_matches_re(s):
    matches = set(pattern.findall(s))
    if matches:
        return True, ", ".join(sorted(matches))
    return False, pd.NA

df[["inlist", "str_found"]] = df["only_col"].apply(
    lambda x: pd.Series(find_matches_re(x))
)

原参考代码问题说明

你提供的参考代码使用set(fruit_list).intersection(set(col_list)),这是完全匹配逻辑,仅当DataFrame元素与fruit_list元素完全相等时才会匹配,无法实现需求中的部分子串匹配,因此不符合要求。

内容的提问来源于stack exchange,提问作者asd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 11:15:41