You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中计算两段文本BLEU分数得分为0的问题排查

排查NLTK sentence_bleu返回0分的问题

问题原因

  1. 默认BLEU-4的计算逻辑:sentence_bleu默认计算BLEU-4,即对1-4元语法(n-gram)的精度取几何平均,且四个n-gram的权重各占25%。只要其中任意一个n-gram的匹配精度为0,最终分数就会直接变为0。
    你的预测文本和参考文本差异较大,比如completed my bachelor's degree vs finished my four-year certification、computer application vs PC application,两者没有任何重合的4元语法,导致4-gram精度为0,直接拉低整体分数至0。
  2. 分词方式粗糙:用split()分词会保留标点与单词的连接(如ABC.),且无法处理缩写(参考里的I'm被当作单个token,而预测里是I和am两个独立token),进一步减少了可匹配的n-gram数量。

解决方法

1. 调整n-gram权重,计算低阶BLEU分数

如果不需要高阶n-gram的匹配,可以调整权重参数,只计算BLEU-1(单字匹配)或BLEU-2(双字匹配):

from nltk.translate.bleu_score import sentence_bleu

prediction = "I am ABC. I have completed my bachelor's degree in computer application at XYZ University and I am currently pursuing my master's degree in computer application through distance education."
reference = "I'm ABC. I have finished my four-year certification in PC application at XYZ and I'm currently pursuing my graduate degree in PC application through distance training."

prediction_tokens = prediction.split()
reference_tokens = reference.split()

# 计算BLEU-1(仅单字匹配)
bleu_1 = sentence_bleu([reference_tokens], prediction_tokens, weights=(1, 0, 0, 0))
print(f"BLEU-1 score: {bleu_1:.4f}")

# 计算BLEU-2(单字+双字匹配)
bleu_2 = sentence_bleu([reference_tokens], prediction_tokens, weights=(0.5, 0.5, 0, 0))
print(f"BLEU-2 score: {bleu_2:.4f}")

2. 使用平滑方法规避高n-gram精度为0的问题

NLTK内置了多种平滑函数,能避免因高阶n-gram无匹配导致的0分:

from nltk.translate.bleu_score import sentence_bleu, SmoothingFunction

prediction_tokens = prediction.split()
reference_tokens = reference.split()

# 使用method4平滑函数
smoothie = SmoothingFunction().method4
smoothed_bleu = sentence_bleu([reference_tokens], prediction_tokens, smoothing_function=smoothie)
print(f"Smoothed BLEU score: {smoothed_bleu:.4f}")

3. 优化分词逻辑

用NLTK的word_tokenize替代split(),能更好处理缩写、标点,同时统一大小写可提升匹配率:

from nltk.tokenize import word_tokenize
from nltk.translate.bleu_score import sentence_bleu, SmoothingFunction

prediction_tokens = word_tokenize(prediction.lower())
reference_tokens = word_tokenize(reference.lower())

smoothie = SmoothingFunction().method4
optimized_bleu = sentence_bleu([reference_tokens], prediction_tokens, smoothing_function=smoothie)
print(f"Optimized BLEU score: {optimized_bleu:.4f}")

内容的提问来源于stack exchange,提问作者Zikra Noman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 06:24:50