使用sklearn计算precision和recall时触发ValueError求助
解决sklearn precision_score触发ValueError的问题
嘿,我来帮你排查这个报错的问题!
首先看你的报错信息:
ValueError: pos_label=None is not a valid label: array(['negative', 'positive'], dtype='<U8')
问题根源
sklearn的precision_score默认情况下,正类标签(pos_label)是数值1,但你的标签是字符串类型的'positive'和'negative',没有数值1这个标签,所以默认参数完全不匹配,直接触发了错误。
顺带提一句,你的预测结果全是'positive',这不会影响报错的解决,先把参数匹配的问题搞定就行。
解决方法
有两种简单的方式可以解决这个问题:
方法1:显式指定正类标签
直接在调用precision_score和recall_score时,通过pos_label参数告诉模型哪个是正类,比如你要把'positive'作为正类的话:
from sklearn.metrics import precision_score, recall_score # 指定pos_label为'positive' precision = precision_score(y_test, pred, pos_label='positive') recall = recall_score(y_test, pred, pos_label='positive')
方法2:把标签转换成数值类型
如果你习惯用数值标签,也可以用LabelEncoder把字符串标签转成0和1,这样就能直接用默认参数了:
from sklearn.metrics import precision_score, recall_score from sklearn.preprocessing import LabelEncoder # 初始化编码器并转换标签 le = LabelEncoder() y_test_encoded = le.fit_transform(y_test) pred_encoded = le.transform(pred) # 现在可以直接调用,默认pos_label=1对应原来的'positive' precision = precision_score(y_test_encoded, pred_encoded) recall = recall_score(y_test_encoded, pred_encoded)
额外说明
等报错解决后,结合你的预测结果,计算出来的指标会有对应的含义:
- precision:所有预测为正类的样本里,真正是正类的比例
- recall:所有真实正类的样本里,被正确预测为正类的比例(这里应该是100%,因为所有正类都被预测成正类了)
内容的提问来源于stack exchange,提问作者merklexy
相关产品推荐
相关产品推荐

