Python如何实现Pandas DataFrame列与指定字符串列表的精准匹配
关键词匹配优化方案
完整实现代码
import pandas as pd import re # 自定义关键词列表 the_list = ['AI', 'NLP', 'approach', 'AR Cloud', 'Army_Intelligence', 'Artificial general intelligence', 'Artificial tissue', 'artificial_insemination', 'artificial_intelligence', 'augmented intelligence', 'augmented reality', 'authentification', 'automaton', 'Autonomous driving', 'Autonomous vehicles', 'bidirectional brain-machine interfaces', 'Biodegradable', 'biodegradable', 'Biotech', 'biotech', 'biotechnology', 'BMI', 'BMIs', 'body_mass_index', 'bourdon', 'Bradypus_tridactylus', 'cognitive computing', 'commercial UAVs', 'Composite AI', 'connected home', 'conversational systems', 'conversational user interfaces', 'dawdler', 'Decentralized web', 'Deep fakes', 'Deep learning', 'defrayal'] # 构造正则匹配规则 # 1. 对每个关键词做转义处理,避免关键词内的正则特殊字符影响匹配 # 2. 用\b单词边界包裹关键词,确保仅匹配独立完整的关键词,避免子串误匹配 pattern = r'\b(' + '|'.join(re.escape(keyword) for keyword in the_list) + r')\b' # 合并标题和内容两个文本列,同时匹配两处的关键词 df['full_text'] = df['title_lemmatized'] + ' ' + df['text_lemmatized'] # 用findall提取所有匹配结果,开启大小写不敏感匹配 df['matched_word(s)'] = df['full_text'].str.findall(pattern, flags=re.IGNORECASE) # (可选)如果需要将匹配结果转为逗号分隔的字符串格式,而非列表格式,可执行下面这行 df['matched_word(s)'] = df['matched_word(s)'].apply(lambda x: ', '.join(x) if x else '') # (可选)如果需要过滤掉无匹配的行,可执行下面这行 # df = df[df['matched_word(s)'] != '']
修改说明
- 解决子串误匹配问题:通过
\b单词边界限定匹配范围,只有关键词作为独立词汇出现时才会被命中,不会匹配其他单词内部的子串 - 解决多匹配仅返回第一个的问题:用
str.findall()替代原有的str.extract(),会返回该行所有符合规则的匹配结果 - 适配双文本列匹配需求:提前将标题、内容两列合并为一个文本字段后再做匹配,确保两处的关键词都能被识别
- 兼容大小写差异:保留
re.IGNORECASE参数,可匹配文本中任意大小写格式的关键词
内容的提问来源于stack exchange,提问作者Zion
相关产品推荐
相关产品推荐

