You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何实现Pandas DataFrame列与指定字符串列表的精准匹配

关键词匹配优化方案

完整实现代码

import pandas as pd
import re

# 自定义关键词列表
the_list = ['AI', 'NLP', 'approach', 'AR Cloud', 'Army_Intelligence', 'Artificial general intelligence', 'Artificial tissue', 'artificial_insemination', 'artificial_intelligence', 'augmented intelligence', 'augmented reality', 'authentification', 'automaton', 'Autonomous driving', 'Autonomous vehicles', 'bidirectional brain-machine interfaces', 'Biodegradable', 'biodegradable', 'Biotech', 'biotech', 'biotechnology', 'BMI', 'BMIs', 'body_mass_index', 'bourdon', 'Bradypus_tridactylus', 'cognitive computing', 'commercial UAVs', 'Composite AI', 'connected home', 'conversational systems', 'conversational user interfaces', 'dawdler', 'Decentralized web', 'Deep fakes', 'Deep learning', 'defrayal']

# 构造正则匹配规则
# 1. 对每个关键词做转义处理,避免关键词内的正则特殊字符影响匹配
# 2. 用\b单词边界包裹关键词,确保仅匹配独立完整的关键词,避免子串误匹配
pattern = r'\b(' + '|'.join(re.escape(keyword) for keyword in the_list) + r')\b'

# 合并标题和内容两个文本列,同时匹配两处的关键词
df['full_text'] = df['title_lemmatized'] + ' ' + df['text_lemmatized']

# 用findall提取所有匹配结果,开启大小写不敏感匹配
df['matched_word(s)'] = df['full_text'].str.findall(pattern, flags=re.IGNORECASE)

# (可选)如果需要将匹配结果转为逗号分隔的字符串格式,而非列表格式,可执行下面这行
df['matched_word(s)'] = df['matched_word(s)'].apply(lambda x: ', '.join(x) if x else '')

# (可选)如果需要过滤掉无匹配的行,可执行下面这行
# df = df[df['matched_word(s)'] != '']

修改说明

  • 解决子串误匹配问题:通过\b单词边界限定匹配范围,只有关键词作为独立词汇出现时才会被命中,不会匹配其他单词内部的子串
  • 解决多匹配仅返回第一个的问题:用str.findall()替代原有的str.extract(),会返回该行所有符合规则的匹配结果
  • 适配双文本列匹配需求:提前将标题、内容两列合并为一个文本字段后再做匹配,确保两处的关键词都能被识别
  • 兼容大小写差异:保留re.IGNORECASE参数,可匹配文本中任意大小写格式的关键词

内容的提问来源于stack exchange,提问作者Zion

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 09:36:07