如何用pandas筛选包含指定字符串列表元素的行并新增匹配字段
Pandas 实现代码
import pandas as pd import re # 1. 构造示例数据 df = pd.DataFrame({ 'id': [1, 2, 3, 4], 'value': [ 'The grapefruit is delicious! But the pear tastes awful.', 'I am a big fan of apple products', 'The quick brown fox jumps over the lazy dog', 'An apple a day keeps the doctor away' ] }) # 2. 定义待匹配的字符串列表 strings = ['apple', 'pear', 'grapefruit'] # 3. 构造正则匹配模式,用re.escape避免字符串中的正则特殊字符影响匹配结果 pattern = '|'.join(re.escape(s) for s in strings) # 4. 提取每行所有匹配的子串,拼接为逗号分隔的字符串 df['value contains substrings:'] = df['value'].str.findall(pattern).apply(', '.join) # 5. 筛选出有匹配结果的行,重置索引 result = df[df['value contains substrings:'] != ''].reset_index(drop=True) print(result)
代码说明
str.findall(pattern)会返回每行中所有符合匹配规则的子串组成的列表,没有匹配就返回空列表。如果需要忽略大小写匹配,可以增加参数flags=re.IGNORECASEapply(', '.join)把匹配到的列表转为逗号分隔的字符串格式,和你需要的输出结构对齐- 最后一步过滤掉没有任何匹配的行,就得到最终结果
运行上述代码输出的result和你给出的目标结构完全一致。
内容的提问来源于stack exchange,提问作者NAShern
相关产品推荐
相关产品推荐

