TensorFlow中测试准确率与混淆矩阵结果不一致求助
问题:自定义数据生成器导致模型评估准确率与混淆矩阵结果不符
更新:问题根源锁定在使用自定义生成器的predict函数上。一次性加载所有测试数据预测时,混淆矩阵准确率符合预期(约90%),但如果无法一次性获取全部测试数据,找不到小批量生成器失效的原因,需要帮助。
我基于COVID-19 Radiography Dataset构建二分类图像分类器,区分COVID感染肺部图像和正常肺部图像。为了利用数据集附带的肺部掩码,我写了一个自定义数据生成器,以批量方式将肺部图像和掩码逐元素相乘处理。
代码运行看似正常,但出现了矛盾结果:
- 模型
evaluate输出的测试准确率约91% - 用scikit-learn生成的混淆矩阵、手动计算的准确率仅约62%
- 手动抽查100个预测结果,错误数约9个,和
evaluate的准确率一致
排查过生成器的shuffle问题,即使全程关闭shuffle,结果依然相同(当前设置训练时打乱,测试时不打乱)。
自定义数据生成器代码
def custom_generator(image_paths, mask_paths, labels, batch_size=32, shuffle=True): num_samples = len(image_paths) while True: # Keras生成器需要无限循环 indices = np.arange(num_samples) if shuffle: np.random.shuffle(indices) for start in range(0, num_samples, batch_size): end = min(start + batch_size, num_samples) # 正确处理最后一批的索引 batch_indices = indices[start:end] # 选择当前批次的索引 # 诊断用打印(已注释) """ print(f"Generating batch: {start // batch_size + 1}") print(f"Start index: {start}, End index: {end}") print(f"Number of indices in this batch: {len(batch_indices)}") """ batch_images = [] batch_labels = [] for idx in batch_indices: img_path = image_paths[int(idx)] mask_path = mask_paths[int(idx)] # 加载并处理图像和掩码 image = Image.open(img_path).convert("L") mask = Image.open(mask_path).convert("L") mask = mask.resize((299, 299), Image.LANCZOS) image_array = np.array(image) / 255.0 mask_array = np.array(mask) / 255.0 multiplied_image = image_array * mask_array # 扩展维度添加通道 batch_images.append(np.expand_dims(multiplied_image, axis=-1)) batch_labels.append(labels[int(idx)]) yield np.array(batch_images), np.array(batch_labels)
模型评估与混淆矩阵代码
# 初始化测试生成器 test_generator = custom_generator(test_image_paths, test_mask_paths, test_labels, batch_size=1, shuffle=False) test_steps = len(test_image_paths) test_loss, test_accuracy = model.evaluate(test_generator, steps=test_steps) print(f"Test Accuracy: {test_accuracy * 100:.2f}%") print(f"Test Loss: {test_loss:.4f}") # 生成预测结果 test_predictions = model.predict(test_generator, steps=test_steps) predicted_classes = (test_predictions > 0.5).astype("int32") # 重新初始化生成器(这里有问题) test_generator = custom_generator(test_image_paths, test_mask_paths, test_labels, batch_size=1) # 计算混淆矩阵和准确率 from sklearn.metrics import confusion_matrix, accuracy_score cm = confusion_matrix(test_labels, predicted_classes) print("Confusion Matrix:") print(cm) cm_accuracy = np.trace(cm) / np.sum(cm) print(f"Accuracy from Confusion Matrix: {cm_accuracy * 100:.2f}%") manual_accuracy = accuracy_score(test_labels, predicted_classes) print(f"Manually Calculated Accuracy: {manual_accuracy * 100:.2f}%")
问题分析与解决方案
核心问题:生成器的状态不一致
Keras的生成器是有状态的迭代器,model.evaluate会耗尽生成器的迭代状态,之后直接调用model.predict时,生成器会从上次结束的位置继续迭代,而不是从头开始。这就导致predicted_classes的顺序和test_labels完全不匹配,混淆矩阵自然错误。
另外,你后续还重新初始化了一个带shuffle=True的生成器,但这步完全多余,反而加剧了混乱。
修复步骤
- 避免重复使用同一个生成器实例:每次调用
evaluate或predict前,都要重新初始化生成器,确保从数据集开头开始迭代 - 统一生成器的shuffle设置:测试阶段必须关闭shuffle,保证预测结果和标签的顺序严格对应
修改后的评估代码:
# 评估模型:重新初始化生成器 test_generator_eval = custom_generator(test_image_paths, test_mask_paths, test_labels, batch_size=1, shuffle=False) test_steps = len(test_image_paths) test_loss, test_accuracy = model.evaluate(test_generator_eval, steps=test_steps) print(f"Test Accuracy: {test_accuracy * 100:.2f}%") print(f"Test Loss: {test_loss:.4f}") # 生成预测:再次初始化全新的生成器 test_generator_pred = custom_generator(test_image_paths, test_mask_paths, test_labels, batch_size=1, shuffle=False) test_predictions = model.predict(test_generator_pred, steps=test_steps) predicted_classes = (test_predictions > 0.5).astype("int32") # 计算混淆矩阵和准确率 from sklearn.metrics import confusion_matrix, accuracy_score cm = confusion_matrix(test_labels, predicted_classes) print("Confusion Matrix:") print(cm) cm_accuracy = np.trace(cm) / np.sum(cm) print(f"Accuracy from Confusion Matrix: {cm_accuracy * 100:.2f}%") manual_accuracy = accuracy_score(test_labels, predicted_classes) print(f"Manually Calculated Accuracy: {manual_accuracy * 100:.2f}%")
针对无法一次性加载数据的场景优化
如果测试集太大无法一次性加载,可以:
- 在生成器迭代时同步收集真实标签,确保预测和标签的顺序完全对应
- 或者将生成器的输出缓存到磁盘(如分批次保存为.npy文件),再批量处理预测结果
示例:同步收集标签的方式
test_generator = custom_generator(test_image_paths, test_mask_paths, test_labels, batch_size=1, shuffle=False) # 同时收集预测结果和真实标签 predicted_classes = [] true_labels = [] for _ in range(test_steps): imgs, lbls = next(test_generator) pred = model.predict(imgs, verbose=0) predicted_classes.append((pred > 0.5).astype("int32")[0][0]) true_labels.append(lbls[0][0]) # 现在计算混淆矩阵 cm = confusion_matrix(true_labels, predicted_classes) # ... 后续计算准确率
这样即使生成器有状态问题,也能保证预测和标签一一对应,不会出现顺序错位。
内容的提问来源于stack exchange,提问作者Dawson Kimyon
相关产品推荐
相关产品推荐

