Python中计算两段文本BLEU分数得分为0的问题排查
排查NLTK sentence_bleu返回0分的问题
问题原因
- 默认BLEU-4的计算逻辑:
sentence_bleu默认计算BLEU-4,即对1-4元语法(n-gram)的精度取几何平均,且四个n-gram的权重各占25%。只要其中任意一个n-gram的匹配精度为0,最终分数就会直接变为0。
你的预测文本和参考文本差异较大,比如completed my bachelor's degreevsfinished my four-year certification、computer applicationvsPC application,两者没有任何重合的4元语法,导致4-gram精度为0,直接拉低整体分数至0。 - 分词方式粗糙:用
split()分词会保留标点与单词的连接(如ABC.),且无法处理缩写(参考里的I'm被当作单个token,而预测里是I和am两个独立token),进一步减少了可匹配的n-gram数量。
解决方法
1. 调整n-gram权重,计算低阶BLEU分数
如果不需要高阶n-gram的匹配,可以调整权重参数,只计算BLEU-1(单字匹配)或BLEU-2(双字匹配):
from nltk.translate.bleu_score import sentence_bleu prediction = "I am ABC. I have completed my bachelor's degree in computer application at XYZ University and I am currently pursuing my master's degree in computer application through distance education." reference = "I'm ABC. I have finished my four-year certification in PC application at XYZ and I'm currently pursuing my graduate degree in PC application through distance training." prediction_tokens = prediction.split() reference_tokens = reference.split() # 计算BLEU-1(仅单字匹配) bleu_1 = sentence_bleu([reference_tokens], prediction_tokens, weights=(1, 0, 0, 0)) print(f"BLEU-1 score: {bleu_1:.4f}") # 计算BLEU-2(单字+双字匹配) bleu_2 = sentence_bleu([reference_tokens], prediction_tokens, weights=(0.5, 0.5, 0, 0)) print(f"BLEU-2 score: {bleu_2:.4f}")
2. 使用平滑方法规避高n-gram精度为0的问题
NLTK内置了多种平滑函数,能避免因高阶n-gram无匹配导致的0分:
from nltk.translate.bleu_score import sentence_bleu, SmoothingFunction prediction_tokens = prediction.split() reference_tokens = reference.split() # 使用method4平滑函数 smoothie = SmoothingFunction().method4 smoothed_bleu = sentence_bleu([reference_tokens], prediction_tokens, smoothing_function=smoothie) print(f"Smoothed BLEU score: {smoothed_bleu:.4f}")
3. 优化分词逻辑
用NLTK的word_tokenize替代split(),能更好处理缩写、标点,同时统一大小写可提升匹配率:
from nltk.tokenize import word_tokenize from nltk.translate.bleu_score import sentence_bleu, SmoothingFunction prediction_tokens = word_tokenize(prediction.lower()) reference_tokens = word_tokenize(reference.lower()) smoothie = SmoothingFunction().method4 optimized_bleu = sentence_bleu([reference_tokens], prediction_tokens, smoothing_function=smoothie) print(f"Optimized BLEU score: {optimized_bleu:.4f}")
内容的提问来源于stack exchange,提问作者Zikra Noman
相关产品推荐
相关产品推荐

