百万行Pandas单列高效匹配水果列表(部分匹配)方案问询
高效实现百万行DataFrame的多子串部分匹配
核心需求
处理百万行单列DataFrame,判断每行内容是否包含fruit_list中的任意子串,同时收集所有匹配的子串,要求性能优于启用regex=True的str.contains方法。
最优方案:Aho-Corasick多模式匹配(pyahocorasick库)
Aho-Corasick算法专门用于多子串批量匹配,预编译所有子串后,每个字符串只需遍历一次,时间复杂度接近线性,适合百万级数据场景。
步骤与代码
- 安装依赖库
pip install pyahocorasick
- 实现匹配逻辑
import pandas as pd import ahocorasick # 示例数据 fruit_list = ["app", "appl", "banana", "pear"] df = pd.DataFrame({ "only_col": ["apple", "banana", "cherry"] }) # 构建AC自动机并加载所有子串 auto = ahocorasick.Automaton() for s in fruit_list: auto.add_word(s, s) auto.make_automaton() # 定义匹配函数:返回是否匹配、匹配的子串列表 def find_matches(s): matches = set() for _, match in auto.iter(s): matches.add(match) if matches: return True, ", ".join(sorted(matches)) return False, pd.NA # 应用到DataFrame,生成目标列 df[["inlist", "str_found"]] = df["only_col"].apply( lambda x: pd.Series(find_matches(x)) )
- 输出结果
only_col inlist str_found 0 apple True app, appl 1 banana True banana 2 cherry False <NA>
纯Pandas替代方案(无需第三方库)
如果不想安装外部库,可预编译优化后的正则表达式,性能略逊于AC自动机,但优于原生循环:
import pandas as pd import re fruit_list = ["app", "appl", "banana", "pear"] df = pd.DataFrame({ "only_col": ["apple", "banana", "cherry"] }) # 预编译正则(转义特殊字符避免匹配异常) pattern = re.compile("|".join(re.escape(s) for s in fruit_list)) def find_matches_re(s): matches = set(pattern.findall(s)) if matches: return True, ", ".join(sorted(matches)) return False, pd.NA df[["inlist", "str_found"]] = df["only_col"].apply( lambda x: pd.Series(find_matches_re(x)) )
原参考代码问题说明
你提供的参考代码使用set(fruit_list).intersection(set(col_list)),这是完全匹配逻辑,仅当DataFrame元素与fruit_list元素完全相等时才会匹配,无法实现需求中的部分子串匹配,因此不符合要求。
内容的提问来源于stack exchange,提问作者asd
相关产品推荐
相关产品推荐

