二分类邮件模型测试预测输出重复数组,如何生成单样本每行一个得分的文件
模型预测得分输出格式异常修复方案
数据集结构
train: 2000 files hit.txt nohit.txt hit.txt Test: 1500 hit.txt nohit.txt hit.txt
原始问题代码
test_dir = 'test/' dictionary = make_dic(test_dir) features_, labels_ = make_dataset(dictionary) calibrated_pred_final = calibrated_clf_pipe.predict(features_) # 输出:array([1, 1, 1, ..., 1, 1, 0]) test_pred_final = calibrated_clf_pipe.predict_proba(features_) import numpy as np batch_y = np.array(test_pred_final).flatten() f = open('scores.txt', 'w') for i in range(len(batch_y)): f.write(str(batch_y))
错误表现
生成的scores.txt为完整数组重复写入的格式:
[0.38636364 0.61363636 0.05147059 ... 0.61363636 0.86734694 0.13265306][0.38636364 0.61363636 0.05147059 ... 0.61363636 0.86734694 0.13265306]
预期输出
每个测试样本对应一个得分,每行输出一个:
0.38636364 0.61363636 0.05147059 ...
问题原因
- 写入逻辑错误:遍历数组下标时,每次写入的是完整
batch_y数组的字符串形式,循环执行次数等于数组长度,导致数组整体被重复写入多次。 - 二分类场景下
predict_proba输出处理不规范:predict_proba返回维度为(样本数, 类别数)的数组,直接全量flatten会把每个样本的两类得分都展开,若仅需正类得分应单独提取对应列。
修复代码
通用修复(保留原有flatten逻辑)
替换写入部分代码即可,使用with上下文管理器自动管理文件句柄:
with open('scores.txt', 'w') as f: for score in batch_y: f.write(f"{score}\n")
二分类场景优化方案(仅输出正类得分)
import numpy as np test_pred_final = calibrated_clf_pipe.predict_proba(features_) # 提取所有样本的正类(第1类)得分,维度为(样本数,)无需额外flatten batch_y = test_pred_final[:, 1] # 按行写入文件,可自定义保留小数位数 with open('scores.txt', 'w') as f: for score in batch_y: f.write(f"{score:.8f}\n")
内容的提问来源于stack exchange,提问作者KSp
相关产品推荐
相关产品推荐

