如何基于字典匹配DataFrame列字符串首个关键词生成新列
实现方案
提供两种常用实现方案,可根据实际数据量和匹配需求选择:
方案1:正则提取+字典映射(推荐,性能更高)
适合数据量较大的场景,通过预编译正则批量提取第一个匹配的key,再映射为对应value:
import pandas as pd import re # 示例数据构造 dictionary = {'dog':'yellow', 'cat':'black', 'frog':'green', 'horse':'brown'} df = pd.DataFrame({ 'ColA': [ 'The dog and horse ate food', 'Where is the frog?', 'horse and cat and frog walked together' ] }) # 编译正则匹配模式,按key长度倒序排序避免短key优先匹配的问题 pattern = re.compile('|'.join(sorted(dictionary.keys(), key=len, reverse=True))) # 可选优化:需要全单词匹配时加上边界符\b,避免部分匹配错误 # pattern = re.compile(r'\b(' + '|'.join(sorted(dictionary.keys(), key=len, reverse=True)) + r')\b') # 提取第一个匹配的key,映射为字典value赋值给新列 df['ColB'] = df['ColA'].str.extract(f'({pattern.pattern})', expand=False).map(dictionary)
方案2:自定义函数遍历匹配(灵活度更高)
适合需要自定义匹配规则的场景,可灵活调整分词、匹配逻辑:
def get_first_match(text): # 拆分所有单词(自动过滤标点) words = re.findall(r'\w+', text) for word in words: if word in dictionary: return dictionary[word] # 无匹配时返回默认值,可根据需求修改 return None df['ColB'] = df['ColA'].apply(get_first_match)
两种方案运行后得到的结果和预期输出完全一致。
内容的提问来源于stack exchange,提问作者dmd7
相关产品推荐
相关产品推荐

