You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何调整Python代码实现大文本语料中非相邻多关键词匹配?

问题描述

需要在数千份文本组成的大语料中匹配非相邻关键词,匹配成功则分配对应标签,否则分配“unknown”标签。

匹配示例

需匹配关键词组sales representative(含同义表述sales rep/customer rep)和dealt,匹配成功则分配类别keyword pattern A:

文本:"The sales representative dealt with everything. It was very helpful to know that he compiled the best option for me."

当前代码问题

已实现单字/相邻词匹配,但非相邻关键词匹配逻辑出错:

  1. 第一版代码仅检查第一列关键词,导致第3、5条文本被错误打标:
import pandas as pd 

# 创建模拟字典
Dict = pd.DataFrame({'word1':['dealt','dealt','dealt',''],
                     'word2':['sales representative','sales rep', 'customer rep', 'options']
                      }  )

# 创建样本文本
texts = ["The sales representative dealt with everything.",
"The sales rep dealt with everything.",
"The agent answered all questions" ,
"The customer rep answered all questions.",
"The agent dealt with everything."]

motive =[]
# 仅检查第一列关键词
for item in texts:
    item = str(item)
    if any(x in item for x in Dict['word1']):
        motive.append('keyword pattern A')        
    else:
        motive.append('unknown')
  1. 第二版代码逻辑混乱,运行后未分配任何标签:
for item in texts:
    #convert into string
    item = str(item)
    #check if keyword can be found in first column
    tempM1 = {x for x in Dict['word1'] if x in item}
    #check if keyword was found
    if tempM1 != None:
        #if yes, locate all of their positions in the dictionary 
        for i in tempM1:
            i = -1
            #get row index 
            ind = Dict.index[Dict['word1'] == list(tempM1)[i+1]] 
    #gives pandas.core.indexes.base.Index            
    #check if column next to given row index is no empty             
            if pd.isnull(Dict['word2'].iloc[ind]) is False:
                #match keyword in second column
                tempM2 = {x for x in Dict['word2'] if x in item}
                #if second keyword was found
                if tempM2 != None: 
                    motive.append('keyword pattern A')
                else: 
            #check again first keyword column
                    tempM3 = {x for x in Dict['word1'] if x in item}
                    if tempM3 != None:
                        motive.append('keyword pattern A')
                    else: 
                        motive.append('unknown')

限制条件

  • 关键词数量约700-1000,组合复杂,不考虑正则方案(认为代码量大且效率低)
  • 项目要求可解释性和透明度,不考虑深度学习及embeddings方案

调整后的代码方案

核心逻辑:先整理字典中的有效关键词组合(分为「双关键词组合」和「单关键词」两类),对每个文本检查是否满足任意一组匹配条件:

import pandas as pd 

# 创建模拟字典
keyword_dict = pd.DataFrame({
    'word1': ['dealt', 'dealt', 'dealt', ''],
    'word2': ['sales representative', 'sales rep', 'customer rep', 'options']
})

# 创建样本文本
texts = [
    "The sales representative dealt with everything.",
    "The sales rep dealt with everything.",
    "The agent answered all questions",
    "The customer rep answered all questions.",
    "The agent dealt with everything."
]

# 预处理:整理有效关键词组合
# 1. 双关键词组合:word1和word2都非空的行
dual_patterns = keyword_dict[(keyword_dict['word1'] != '') & (keyword_dict['word2'] != '')]
# 2. 单关键词:word1或word2为空的行(提取非空的那个)
single_patterns = []
for _, row in keyword_dict.iterrows():
    if row['word1'] == '' and row['word2'] != '':
        single_patterns.append(row['word2'])
    elif row['word2'] == '' and row['word1'] != '':
        single_patterns.append(row['word1'])

motive = []
for text in texts:
    text_str = str(text)
    matched = False
    
    # 检查双关键词组合:同一行的word1和word2都在文本中(不考虑顺序和中间内容)
    for _, row in dual_patterns.iterrows():
        if row['word1'] in text_str and row['word2'] in text_str:
            matched = True
            break
    
    # 如果双关键词没匹配到,检查单关键词
    if not matched:
        for keyword in single_patterns:
            if keyword in text_str:
                matched = True
                break
    
    # 根据匹配结果打标签
    motive.append('keyword pattern A' if matched else 'unknown')

# 输出结果
print(motive)
# 预期输出:['keyword pattern A', 'keyword pattern A', 'unknown', 'keyword pattern A', 'unknown']

代码说明

  1. 预处理字典:把原字典拆分为「双关键词组合」和「单关键词」,避免重复检查
  2. 双关键词匹配:确保同一行的两个关键词都出现在文本中(不关心位置和中间内容)
  3. 单关键词匹配:仅当双关键词没匹配到的时候,检查是否有单关键词存在
  4. 逻辑清晰:每一步都有明确的判断,可解释性强,且效率优于原代码

内容的提问来源于stack exchange,提问作者Simone

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 02:33:32