如何调整Python代码实现大文本语料中非相邻多关键词匹配?
问题描述
需要在数千份文本组成的大语料中匹配非相邻关键词,匹配成功则分配对应标签,否则分配“unknown”标签。
匹配示例
需匹配关键词组sales representative(含同义表述sales rep/customer rep)和dealt,匹配成功则分配类别keyword pattern A:
文本:"The sales representative dealt with everything. It was very helpful to know that he compiled the best option for me."
当前代码问题
已实现单字/相邻词匹配,但非相邻关键词匹配逻辑出错:
- 第一版代码仅检查第一列关键词,导致第3、5条文本被错误打标:
import pandas as pd # 创建模拟字典 Dict = pd.DataFrame({'word1':['dealt','dealt','dealt',''], 'word2':['sales representative','sales rep', 'customer rep', 'options'] } ) # 创建样本文本 texts = ["The sales representative dealt with everything.", "The sales rep dealt with everything.", "The agent answered all questions" , "The customer rep answered all questions.", "The agent dealt with everything."] motive =[] # 仅检查第一列关键词 for item in texts: item = str(item) if any(x in item for x in Dict['word1']): motive.append('keyword pattern A') else: motive.append('unknown')
- 第二版代码逻辑混乱,运行后未分配任何标签:
for item in texts: #convert into string item = str(item) #check if keyword can be found in first column tempM1 = {x for x in Dict['word1'] if x in item} #check if keyword was found if tempM1 != None: #if yes, locate all of their positions in the dictionary for i in tempM1: i = -1 #get row index ind = Dict.index[Dict['word1'] == list(tempM1)[i+1]] #gives pandas.core.indexes.base.Index #check if column next to given row index is no empty if pd.isnull(Dict['word2'].iloc[ind]) is False: #match keyword in second column tempM2 = {x for x in Dict['word2'] if x in item} #if second keyword was found if tempM2 != None: motive.append('keyword pattern A') else: #check again first keyword column tempM3 = {x for x in Dict['word1'] if x in item} if tempM3 != None: motive.append('keyword pattern A') else: motive.append('unknown')
限制条件
- 关键词数量约700-1000,组合复杂,不考虑正则方案(认为代码量大且效率低)
- 项目要求可解释性和透明度,不考虑深度学习及embeddings方案
调整后的代码方案
核心逻辑:先整理字典中的有效关键词组合(分为「双关键词组合」和「单关键词」两类),对每个文本检查是否满足任意一组匹配条件:
import pandas as pd # 创建模拟字典 keyword_dict = pd.DataFrame({ 'word1': ['dealt', 'dealt', 'dealt', ''], 'word2': ['sales representative', 'sales rep', 'customer rep', 'options'] }) # 创建样本文本 texts = [ "The sales representative dealt with everything.", "The sales rep dealt with everything.", "The agent answered all questions", "The customer rep answered all questions.", "The agent dealt with everything." ] # 预处理:整理有效关键词组合 # 1. 双关键词组合:word1和word2都非空的行 dual_patterns = keyword_dict[(keyword_dict['word1'] != '') & (keyword_dict['word2'] != '')] # 2. 单关键词:word1或word2为空的行(提取非空的那个) single_patterns = [] for _, row in keyword_dict.iterrows(): if row['word1'] == '' and row['word2'] != '': single_patterns.append(row['word2']) elif row['word2'] == '' and row['word1'] != '': single_patterns.append(row['word1']) motive = [] for text in texts: text_str = str(text) matched = False # 检查双关键词组合:同一行的word1和word2都在文本中(不考虑顺序和中间内容) for _, row in dual_patterns.iterrows(): if row['word1'] in text_str and row['word2'] in text_str: matched = True break # 如果双关键词没匹配到,检查单关键词 if not matched: for keyword in single_patterns: if keyword in text_str: matched = True break # 根据匹配结果打标签 motive.append('keyword pattern A' if matched else 'unknown') # 输出结果 print(motive) # 预期输出:['keyword pattern A', 'keyword pattern A', 'unknown', 'keyword pattern A', 'unknown']
代码说明
- 预处理字典:把原字典拆分为「双关键词组合」和「单关键词」,避免重复检查
- 双关键词匹配:确保同一行的两个关键词都出现在文本中(不关心位置和中间内容)
- 单关键词匹配:仅当双关键词没匹配到的时候,检查是否有单关键词存在
- 逻辑清晰:每一步都有明确的判断,可解释性强,且效率优于原代码
内容的提问来源于stack exchange,提问作者Simone
相关产品推荐
相关产品推荐

