You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何判断DataFrame某列是否包含多个列表元素并提取匹配值?

pandas按关键词列表匹配文本生成新列

需求描述

现有3个预设关键词列表,需要在DataFrame的desxription列中检索匹配各列表内的元素:

  • 为每个列表新增和列表同名的新列
  • 若当前行的文本包含对应列表的元素,将匹配到的元素写入该列
  • 无匹配项则填充NA

示例数据

初始DataFrame

id      desxription
1         'this is bad'
2         'city tehran country iran'
3         'uA is a country'
5         'this is summer'
6         'this is winter'
7         'this is canada'
8         'this is toronto'

待匹配关键词列表

L1 = ['summer', 'winter', 'fall']
L2 = ['iran', 'uA']
L3 = ['tehran', 'canada', 'toronto']

预期输出结果

id      desxription                       L1          L2         L3
1         'this is bad'                      NA         NA          NA
2         'city tehran country iran'       NA         iran        tehran
3         'uA is a country'                  NA         uA          NA
5         'this is summer'                  summer      NA          NA
6         'this is winter'                  winter      NA          NA
7         'this is canada'                 NA         NA          canada
8         'this is toronto'               NA         NA          toronto

实现代码

import pandas as pd
import numpy as np
import re

# 构造示例DataFrame
df = pd.DataFrame({
    'id': [1, 2, 3, 5, 6, 7, 8],
    'desxription': [
        'this is bad',
        'city tehran country iran',
        'uA is a country',
        'this is summer',
        'this is winter',
        'this is canada',
        'this is toronto'
    ]
})

# 汇总所有待匹配的列表,key为新列名,value为对应关键词列表
match_config = {
    'L1': L1,
    'L2': L2,
    'L3': L3
}

# 批量匹配生成新列
for col, keywords in match_config.items():
    # 拼接正则规则:\b为单词边界,避免子串误匹配;re.escape处理关键词里的特殊正则字符
    pattern = rf'\b({"|".join(map(re.escape, keywords))})\b'
    # 提取匹配到的关键词,无匹配自动返回NaN
    df[col] = df['desxription'].str.extract(pattern, expand=False)

print(df)

说明

  • 代码默认做整词匹配,不会出现类似关键词sum匹配到summer的误匹配问题,如果不需要整词匹配,删除规则里的两个\b即可
  • 如果需要不区分大小写匹配,给str.extract传入参数flags=re.IGNORECASE
  • 如果单条文本可能匹配到同个列表的多个关键词,可以把str.extract替换为str.findall,得到所有匹配结果的列表后按需处理

内容的提问来源于stack exchange,提问作者user15649753

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 21:57:21