You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于字典匹配DataFrame列字符串首个关键词生成新列

实现方案

提供两种常用实现方案,可根据实际数据量和匹配需求选择:

方案1:正则提取+字典映射(推荐,性能更高)

适合数据量较大的场景,通过预编译正则批量提取第一个匹配的key,再映射为对应value:

import pandas as pd
import re

# 示例数据构造
dictionary = {'dog':'yellow', 'cat':'black', 'frog':'green', 'horse':'brown'}
df = pd.DataFrame({
    'ColA': [
        'The dog and horse ate food',
        'Where is the frog?',
        'horse and cat and frog walked together'
    ]
})

# 编译正则匹配模式,按key长度倒序排序避免短key优先匹配的问题
pattern = re.compile('|'.join(sorted(dictionary.keys(), key=len, reverse=True)))
# 可选优化:需要全单词匹配时加上边界符\b,避免部分匹配错误
# pattern = re.compile(r'\b(' + '|'.join(sorted(dictionary.keys(), key=len, reverse=True)) + r')\b')

# 提取第一个匹配的key,映射为字典value赋值给新列
df['ColB'] = df['ColA'].str.extract(f'({pattern.pattern})', expand=False).map(dictionary)

方案2:自定义函数遍历匹配(灵活度更高)

适合需要自定义匹配规则的场景,可灵活调整分词、匹配逻辑:

def get_first_match(text):
    # 拆分所有单词(自动过滤标点)
    words = re.findall(r'\w+', text)
    for word in words:
        if word in dictionary:
            return dictionary[word]
    # 无匹配时返回默认值,可根据需求修改
    return None

df['ColB'] = df['ColA'].apply(get_first_match)

两种方案运行后得到的结果和预期输出完全一致。

内容的提问来源于stack exchange,提问作者dmd7

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 10:45:03