pandas中如何通过字典匹配列包含的关键字新增对应映射值列
pandas字段匹配映射实现方案
方法1:自定义函数+apply(适合小数据量,逻辑直观)
核心逻辑是定义匹配函数对每个单元格内容遍历字典键做包含校验,返回对应的颜色值,再批量应用到目标列:
import pandas as pd # 构造原始数据 foo_dt = pd.DataFrame({'var_1': ['filter coffee', 'american cheesecake', 'espresso coffee', 'latte tea'], 'var_2': ['coffee', 'coffee black', 'tea', 'strawberry cheesecake']}) foo_colors = {'coffee': 'brown', 'cheesecake': 'white', 'tea': 'green'} # 定义颜色匹配函数 def get_color(text): for keyword, color in foo_colors.items(): if keyword in text: return color # 无匹配场景可自定义返回值,本例无此场景可省略 return None # 生成新列 foo_dt['color_var_1'] = foo_dt['var_1'].apply(get_color) foo_dt['color_var_2'] = foo_dt['var_2'].apply(get_color)
如果存在单个单元格匹配多个字典键的场景,返回结果由字典的遍历顺序决定,可调整字典键的排序控制匹配优先级。
方法2:正则提取+映射(适合大数据量,性能更优)
使用pandas向量化字符串操作,避免逐行遍历的性能损耗:
import pandas as pd foo_dt = pd.DataFrame({'var_1': ['filter coffee', 'american cheesecake', 'espresso coffee', 'latte tea'], 'var_2': ['coffee', 'coffee black', 'tea', 'strawberry cheesecake']}) foo_colors = {'coffee': 'brown', 'cheesecake': 'white', 'tea': 'green'} # 拼接匹配正则 match_pattern = '|'.join(foo_colors.keys()) # 提取关键字后映射为颜色 foo_dt['color_var_1'] = foo_dt['var_1'].str.extract(f'({match_pattern})', expand=False).map(foo_colors) foo_dt['color_var_2'] = foo_dt['var_2'].str.extract(f'({match_pattern})', expand=False).map(foo_colors)
两种方法输出结果均与你给出的预期结果完全一致。
内容的提问来源于stack exchange,提问作者quant
相关产品推荐
相关产品推荐

