You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提升Rake-Nltp中特定词汇被选为关键词的概率?

问题描述

我用Rake-Nltp处理足球赛事解说文本生成关键词,比如对于语句Player1 puts a cross from right wing..,希望cross成为高得分关键词,但默认评分给出的分数很低。我尝试实现了自定义评分函数,但完全没改变结果——哪怕我知道代码没针对特定场景执行,输出也毫无变化。

我的自定义函数和Rake初始化代码如下:

def custom_scoring_function(word_frequency, word_degree, degree_to_frequency_ratio):
    # Increase score for words that match a certain pattern
    scores = {}
    # Loop through all candidate keywords
    for word in word_frequency.keys():
        # Example custom rule: give higher scores to keywords with more than one word
        if len(word.split()) > 1:
            scores[word] = word_degree[word] * 10
        else:
            scores[word] = (word_degree[word] / degree_to_frequency_ratio[word]) * 3
    return scores
r = Rake(ranking_metric=custom_scoring_function, stopwords=stopwords, max_length=3)
问题诊断与解决

1. 先确认自定义函数是否被执行

首先排查函数是否真的在工作:在函数里加入调试打印,检查候选词是否包含cross,以及函数是否被调用:

def custom_scoring_function(word_frequency, word_degree, degree_to_frequency_ratio):
    print("候选词列表:", list(word_frequency.keys()))  # 调试用
    scores = {}
    # 定义需要加权的足球术语
    football_terms = {"cross", "pass", "shot", "goal", "wing"}
    for word in word_frequency.keys():
        # 优先给目标足球术语高权重
        if word in football_terms:
            scores[word] = word_degree[word] * 10
        elif len(word.split()) > 1:
            scores[word] = word_degree[word] * 5
        else:
            # 其他单词保留基础评分逻辑并适度加权
            scores[word] = degree_to_frequency_ratio[word] * 2
    return scores

如果执行后没看到候选词列表的输出,说明ranking_metric参数未被正确识别:

  • 检查Rake-Nltp版本,旧版本可能不支持该参数,执行pip install --upgrade rake-nltk升级到最新版。
  • 确认参数拼写正确,部分版本可能用scoring_function而非ranking_metric。

2. 确保cross未被过滤

如果候选词列表里没有cross,说明它被当成停用词过滤了:

  • 打印stopwords列表,确认cross不在其中。
  • 若使用自定义停用词,检查是否误将足球术语加入了停用词库。

3. 验证调整后的结果

执行以下代码查看关键词及对应分数:

text = "Player1 puts a cross from right wing.."
r.extract_keywords_from_text(text)
# 输出带分数的排序关键词
print(r.get_ranked_phrases_with_scores())

正常情况下会输出类似[(10.0, 'cross'), (2.0, 'right wing'), ...]的结果,cross的分数会显著高于其他词汇。


内容的提问来源于stack exchange,提问作者Ijk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 10:12:41