Python如何实现ngram优先级高于unigram的文本打分函数
实现思路
- 先将原文本拆分为单词列表,同时创建等长的标记数组,用于记录哪些位置的单词已经被匹配过
- 对词典中的所有词条按包含的单词数量降序排序,确保长度更长的ngram优先被处理,满足优先级要求
- 依次遍历排序后的词条,查找文本中连续匹配的位置,只有对应位置的单词都未被使用时才计入得分,同时标记这些单词为已使用,避免后续短的unigram重复统计
完整实现代码
import pandas as pd def calculate_score(text, dictionary): # 拆分原文本为单词列表 text_words = text.split() text_length = len(text_words) # 标记数组:记录对应位置的单词是否已经被匹配占用 used_flags = [False] * text_length total = 0 # 给词条按单词数降序排序,长ngram优先处理 processed_dict = dictionary.copy() processed_dict['word_count'] = processed_dict['text'].apply(lambda x: len(x.split())) processed_dict = processed_dict.sort_values(by='word_count', ascending=False) for _, row in processed_dict.iterrows(): term = row['text'] term_score = row['score'] term_words = term.split() term_length = len(term_words) # 词条长度超过文本总长度直接跳过 if term_length > text_length: continue # 遍历所有可能的起始匹配位置 for start_idx in range(text_length - term_length + 1): # 检查是否完全匹配且所有位置未被占用 if text_words[start_idx:start_idx+term_length] == term_words \ and all(not used_flags[idx] for idx in range(start_idx, start_idx+term_length)): total += term_score # 标记对应位置为已占用 for idx in range(start_idx, start_idx+term_length): used_flags[idx] = True return total
示例测试
# 示例输入 text = "I would like to reduce carbon emissions" dictionary = pd.DataFrame({ 'text': ["like","reduce","carbon","emissions","reduce carbon emissions"], 'score': [1,-1,-1,-1,1] }) # 调用函数计算得分 print(calculate_score(text, dictionary)) # 输出结果:2
内容的提问来源于stack exchange,提问作者Rollo99
相关产品推荐
相关产品推荐

