You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何匹配pandas DataFrame行与字典键值并新增列存储匹配对应值

解决方案

你可以直接复用已经写好的正则匹配结果,只需要新增一行逻辑提取匹配到的具体文本即可,无需修改原有匹配逻辑:

import pandas as pd

entertainment_dict = {
  "Food": ["McDonald", "Five Guys", "KFC"],
  "Music": ["Taylor Swift", "Jay Z", "One Direction"],
  "TV": ["Big Bang Theory", "Queen of South", "Ted Lasso"]
}

data = {'text':["Kevin Lee has bought a Taylor Swift's CD and eaten at McDonald.", 
                "The best burger in McDonald is cheeze buger.",
                "Kevin Lee is planning to watch the Big Bang Theory and eat at KFC."]}

df = pd.DataFrame(data)

regex = '|'.join(f'(?P<{k}>{"|".join(v)})' for k,v in entertainment_dict.items())
# 先把匹配结果存下来,避免重复计算
matches = df['text'].str.extractall(regex)
# 原有labels生成逻辑不变
df['labels'] = ((matches.notnull().groupby(level=0).max()*entertainment_dict.keys())
                 .apply(lambda r: ','.join([i for i in r if i]) , axis=1)
                )
# 新增words列:提取每行所有非空匹配值,用逗号加空格拼接,无匹配的行填充为空字符串
df['words'] = matches.stack().groupby(level=0).apply(', '.join).fillna('')

运行后得到的结果和你给出的预期输出完全一致。

逻辑说明

  • str.extractall返回的结果第一层索引是原DataFrame的行号,第二层索引是匹配次数,每一列对应正则里的命名分组(也就是字典的key),匹配到值就存对应文本,没匹配到就是NaN
  • stack()会自动剔除所有NaN值,仅保留实际匹配到的文本
  • 按第一层索引(原行号)分组后拼接,就能得到每行所有匹配到的具体词汇

内容的提问来源于stack exchange,提问作者codedancer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 12:48:00