如何用列表字符串动态生成正则表达式在Pandas DataFrame中搜索?
动态构建正则表达式实现Pandas多关键词筛选
问题场景
要从关键词列表动态生成带“或”逻辑的正则表达式,用来筛选Pandas DataFrame中包含任一关键词的行,但原代码运行时出现r is not defined错误,无法正常执行。
错误代码示例
kw_list = ["cod", "i"] kw_regex_string = "\b(" for kw in kw_list: kw_regex_string = kw_regex_string + kw + "|" kw_regex_string = kw_regex_string[:-1] # 移除末尾多余的"|" kw_regex_string = kw_regex_string + ")\b" myregex = r + kw_regex_string # 这里报错:r未定义 texts_df.loc[texts_df["text"].str.contains(myregex, regex=True)]
错误原因
r是Python的原始字符串前缀,不是变量,不能直接和字符串相加;- 普通字符串里的
\b会被解析为退格符,而非正则表达式中的单词边界,必须用原始字符串或双反斜杠转义。
正确实现方法
方法1:简洁拼接关键词(无特殊正则字符时用)
用str.join()直接拼接关键词,配合原始字符串定义正则,避免循环拼接的冗余:
import pandas as pd texts_df = pd.DataFrame({"id":[1,2,3,4], "text":["she loves coding", "he was eating cod", "i do not like fish", "fishing is not for me"]}) kw_list = ["cod", "i"] # 用|连接关键词,包裹在单词边界\b中 kw_pattern = r'\b(' + '|'.join(kw_list) + r')\b' # 执行筛选 result = texts_df.loc[texts_df["text"].str.contains(kw_pattern, regex=True)] print(result)
运行后会正确筛选出第2、3行数据。
方法2:处理含特殊正则字符的关键词
如果关键词包含.、*这类正则元字符,需要先用re.escape()转义,避免正则逻辑出错:
import re import pandas as pd texts_df = pd.DataFrame({"id":[1,2,3,4], "text":["she loves coding", "he was eating cod", "i do not like fish", "fishing is not for me"]}) kw_list = ["cod", "i"] # 转义每个关键词中的特殊正则字符 escaped_keywords = [re.escape(kw) for kw in kw_list] kw_pattern = r'\b(' + '|'.join(escaped_keywords) + r')\b' result = texts_df.loc[texts_df["text"].str.contains(kw_pattern, regex=True)] print(result)
内容的提问来源于stack exchange,提问作者code_to_joy
相关产品推荐
相关产品推荐

