分类任务遇‘二元与未知目标混合’报错,如何忽略未知目标计算指标?
解决分类指标计算中混合二进制与未知目标的报错问题
遇到这个报错太正常了——scikit-learn的分类评估指标(比如混淆矩阵)只认纯数值的标签,而你的y_pred里混了None(属于未知目标类型),自然会触发类型不匹配的错误。咱们只需要把那些没有有效预测的样本过滤掉,只保留有明确0/1预测的样本对,就能正常计算指标了。
具体解决方案
核心思路是同时过滤测试集标签y_test和预测结果y_pred,确保两者只保留y_pred不为None的样本,这样就都是合法的二进制标签了。
你可以在生成y_pred之后,添加这段代码完成过滤:
# 生成过滤掩码:标记哪些y_pred不是None mask = [pred is not None for pred in y_pred] # 同步过滤测试集标签和预测结果 y_test_filtered = y_test[mask] y_pred_filtered = [p for p in y_pred if p is not None]
然后把原来计算指标的代码,全部替换成使用过滤后的数组:
from sklearn.metrics import classification_report, confusion_matrix, accuracy_score print('confusion matrix\n',confusion_matrix(y_test_filtered, y_pred_filtered)) print('classification report\n', classification_report(y_test_filtered, y_pred_filtered)) print('accuracy score', accuracy_score(y_test_filtered, y_pred_filtered)) # 注意:MAE、MSE是回归任务的指标,用在分类场景意义不大,可按需保留 from sklearn import metrics print('Mean Absolute Error:', metrics.mean_absolute_error(y_test_filtered, y_pred_filtered)) print('Mean Squared Error:', metrics.mean_squared_error(y_test_filtered, y_pred_filtered)) print('Root Mean Squared Error:', np.sqrt(metrics.mean_squared_error(y_test_filtered, y_pred_filtered)))
整合后的完整关键代码片段
把修改整合到你的原有代码中,最终关键部分如下:
probs = regressor.predict_proba(X_test) # 生成带None的预测结果 y_pred = [None]*len(y_test) for i in range(len(probs)): if (probs[i][0]>=0.8) & (probs[i][1]<0.8): y_pred[i] = 0 elif (probs[i][0]<0.8) & (probs[i][1]>=0.8): y_pred[i] = 1 # 过滤无效预测样本 mask = [pred is not None for pred in y_pred] y_test_filtered = y_test[mask] y_pred_filtered = [p for p in y_pred if p is not None] # 计算分类指标 from sklearn.metrics import classification_report, confusion_matrix, accuracy_score print('confusion matrix\n',confusion_matrix(y_test_filtered, y_pred_filtered)) print('classification report\n', classification_report(y_test_filtered, y_pred_filtered)) print('accuracy score', accuracy_score(y_test_filtered, y_pred_filtered))
这样修改后,就只会基于那些有明确预测结果的样本来计算指标,不会再触发类型不匹配的错误了。
内容的提问来源于stack exchange,提问作者منى
相关产品推荐
相关产品推荐

