You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何实现ngram优先级高于unigram的文本打分函数

实现思路
  • 先将原文本拆分为单词列表,同时创建等长的标记数组,用于记录哪些位置的单词已经被匹配过
  • 对词典中的所有词条按包含的单词数量降序排序,确保长度更长的ngram优先被处理,满足优先级要求
  • 依次遍历排序后的词条,查找文本中连续匹配的位置,只有对应位置的单词都未被使用时才计入得分,同时标记这些单词为已使用,避免后续短的unigram重复统计
完整实现代码
import pandas as pd

def calculate_score(text, dictionary):
    # 拆分原文本为单词列表
    text_words = text.split()
    text_length = len(text_words)
    # 标记数组:记录对应位置的单词是否已经被匹配占用
    used_flags = [False] * text_length
    total = 0

    # 给词条按单词数降序排序,长ngram优先处理
    processed_dict = dictionary.copy()
    processed_dict['word_count'] = processed_dict['text'].apply(lambda x: len(x.split()))
    processed_dict = processed_dict.sort_values(by='word_count', ascending=False)

    for _, row in processed_dict.iterrows():
        term = row['text']
        term_score = row['score']
        term_words = term.split()
        term_length = len(term_words)

        # 词条长度超过文本总长度直接跳过
        if term_length > text_length:
            continue
        
        # 遍历所有可能的起始匹配位置
        for start_idx in range(text_length - term_length + 1):
            # 检查是否完全匹配且所有位置未被占用
            if text_words[start_idx:start_idx+term_length] == term_words \
                and all(not used_flags[idx] for idx in range(start_idx, start_idx+term_length)):
                total += term_score
                # 标记对应位置为已占用
                for idx in range(start_idx, start_idx+term_length):
                    used_flags[idx] = True
    return total
示例测试
# 示例输入
text = "I would like to reduce carbon emissions"
dictionary = pd.DataFrame({
    'text': ["like","reduce","carbon","emissions","reduce carbon emissions"],
    'score': [1,-1,-1,-1,1]
})

# 调用函数计算得分
print(calculate_score(text, dictionary))
# 输出结果:2

内容的提问来源于stack exchange,提问作者Rollo99

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 09:54:03