如何判断DataFrame某列是否包含多个列表元素并提取匹配值?
pandas按关键词列表匹配文本生成新列
需求描述
现有3个预设关键词列表,需要在DataFrame的desxription列中检索匹配各列表内的元素:
- 为每个列表新增和列表同名的新列
- 若当前行的文本包含对应列表的元素,将匹配到的元素写入该列
- 无匹配项则填充NA
示例数据
初始DataFrame
id desxription 1 'this is bad' 2 'city tehran country iran' 3 'uA is a country' 5 'this is summer' 6 'this is winter' 7 'this is canada' 8 'this is toronto'
待匹配关键词列表
L1 = ['summer', 'winter', 'fall'] L2 = ['iran', 'uA'] L3 = ['tehran', 'canada', 'toronto']
预期输出结果
id desxription L1 L2 L3 1 'this is bad' NA NA NA 2 'city tehran country iran' NA iran tehran 3 'uA is a country' NA uA NA 5 'this is summer' summer NA NA 6 'this is winter' winter NA NA 7 'this is canada' NA NA canada 8 'this is toronto' NA NA toronto
实现代码
import pandas as pd import numpy as np import re # 构造示例DataFrame df = pd.DataFrame({ 'id': [1, 2, 3, 5, 6, 7, 8], 'desxription': [ 'this is bad', 'city tehran country iran', 'uA is a country', 'this is summer', 'this is winter', 'this is canada', 'this is toronto' ] }) # 汇总所有待匹配的列表,key为新列名,value为对应关键词列表 match_config = { 'L1': L1, 'L2': L2, 'L3': L3 } # 批量匹配生成新列 for col, keywords in match_config.items(): # 拼接正则规则:\b为单词边界,避免子串误匹配;re.escape处理关键词里的特殊正则字符 pattern = rf'\b({"|".join(map(re.escape, keywords))})\b' # 提取匹配到的关键词,无匹配自动返回NaN df[col] = df['desxription'].str.extract(pattern, expand=False) print(df)
说明
- 代码默认做整词匹配,不会出现类似关键词
sum匹配到summer的误匹配问题,如果不需要整词匹配,删除规则里的两个\b即可 - 如果需要不区分大小写匹配,给
str.extract传入参数flags=re.IGNORECASE - 如果单条文本可能匹配到同个列表的多个关键词,可以把
str.extract替换为str.findall,得到所有匹配结果的列表后按需处理
内容的提问来源于stack exchange,提问作者user15649753
相关产品推荐
相关产品推荐

