You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pandas中如何通过字典匹配列包含的关键字新增对应映射值列

pandas字段匹配映射实现方案

方法1:自定义函数+apply(适合小数据量,逻辑直观)

核心逻辑是定义匹配函数对每个单元格内容遍历字典键做包含校验,返回对应的颜色值,再批量应用到目标列:

import pandas as pd

# 构造原始数据
foo_dt = pd.DataFrame({'var_1': ['filter coffee', 'american cheesecake', 'espresso coffee', 'latte tea'],
                   'var_2': ['coffee', 'coffee black', 'tea', 'strawberry cheesecake']})
foo_colors = {'coffee': 'brown', 'cheesecake': 'white', 'tea': 'green'}

# 定义颜色匹配函数
def get_color(text):
    for keyword, color in foo_colors.items():
        if keyword in text:
            return color
    # 无匹配场景可自定义返回值,本例无此场景可省略
    return None

# 生成新列
foo_dt['color_var_1'] = foo_dt['var_1'].apply(get_color)
foo_dt['color_var_2'] = foo_dt['var_2'].apply(get_color)

如果存在单个单元格匹配多个字典键的场景,返回结果由字典的遍历顺序决定,可调整字典键的排序控制匹配优先级。

方法2:正则提取+映射(适合大数据量,性能更优)

使用pandas向量化字符串操作,避免逐行遍历的性能损耗:

import pandas as pd

foo_dt = pd.DataFrame({'var_1': ['filter coffee', 'american cheesecake', 'espresso coffee', 'latte tea'],
                   'var_2': ['coffee', 'coffee black', 'tea', 'strawberry cheesecake']})
foo_colors = {'coffee': 'brown', 'cheesecake': 'white', 'tea': 'green'}

# 拼接匹配正则
match_pattern = '|'.join(foo_colors.keys())
# 提取关键字后映射为颜色
foo_dt['color_var_1'] = foo_dt['var_1'].str.extract(f'({match_pattern})', expand=False).map(foo_colors)
foo_dt['color_var_2'] = foo_dt['var_2'].str.extract(f'({match_pattern})', expand=False).map(foo_colors)

两种方法输出结果均与你给出的预期结果完全一致。

内容的提问来源于stack exchange,提问作者quant

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 05:06:03