使用NLTK的BLEU对比两个完全相同句子无法得到1.0是什么原因
问题原因
你调用nltk.translate.bleu_score.sentence_bleu时第一个参数的传参格式不符合要求:
该方法的第一个参数需要接收参考译文集合,支持同时传入多条参考译文,每个参考译文本身是分词后的列表,因此入参格式是[reference_1_token_list, reference_2_token_list, ...]。你直接传入了单条分词后的参考列表,NLTK会误将每个独立单词识别为一条单独的参考句子,4-grams匹配完全失效,才会得到极低的异常评分。
修正后的代码
import nltk from nltk import word_tokenize from nltk.translate.bleu_score import SmoothingFunction ref = 'You know that it would be untrue You know that I would be a liar If I was to say to you Girl, we couldnt get much higher.' cand = 'You know that it would be untrue You know that I would be a liar If I was to say to you Girl, we couldnt get much higher.' smoothie = SmoothingFunction().method1 reference = word_tokenize(ref) candidate = word_tokenize(cand) weights = (0.25, 0.25, 0.25, 0.25) # 仅修改第一个参数为 [reference] 即可 BLEUscore = nltk.translate.bleu_score.sentence_bleu([reference], candidate, weights, smoothing_function=smoothie) print(BLEUscore)
运行结果
修正后输出结果为1.0,符合相同句子BLEU评分取满分的预期,你选用的4-grams权重、method1平滑函数适配单句评估场景的设置没有其他问题。
内容的提问来源于stack exchange,提问作者slow_war
相关产品推荐
相关产品推荐

