调用NLTK的meteor_score计算Meteor分数时报类型错误如何解决?
NLTK meteor_score 接口正确调用方案
两次报错的核心原因是不符合接口的参数格式要求:
- 参考句集合和待评估假设句都不能直接传入整句字符串,必须提前拆分为独立分词(token)组成的列表
- 第一个入参是参考句集合,格式为
List[List[str]]:由多个分词后的参考句列表共同组成的二维列表,不能是整句字符串组成的一维列表 - 第二个入参是待评估的假设句,格式为
List[str]:单个句子分词后得到的字符串列表,不能是整句字符串、也不是包裹整句的单层列表
最简实现(手动分词)
适合简单场景,直接手动拆分句子为token:
from nltk.translate.meteor_score import meteor_score # 参考句:每个句子单独分词为列表 references = [ ["this", "is", "an", "apple"], ["that", "is", "an", "apple"] ] # 假设句:同样分词为token列表 hypothesis = ["an", "apple", "on", "this", "tree"] # 计算并输出分数,取值范围0~1,分数越高匹配度越高 print(meteor_score(references, hypothesis))
自动分词实现
如果需要处理大量句子,可以调用NLTK自带的分词工具自动拆分:
from nltk.translate.meteor_score import meteor_score from nltk.tokenize import word_tokenize import nltk # 首次运行需要下载分词依赖包,后续运行可注释该行 nltk.download('punkt') # 自动对整句执行分词 references = [ word_tokenize("this is an apple"), word_tokenize("that is an apple") ] hypothesis = word_tokenize("an apple on this tree") print(meteor_score(references, hypothesis))
内容的提问来源于stack exchange,提问作者ziad Sakr
相关产品推荐
相关产品推荐

