You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow中测试准确率与混淆矩阵结果不一致求助

问题:自定义数据生成器导致模型评估准确率与混淆矩阵结果不符

更新:问题根源锁定在使用自定义生成器的predict函数上。一次性加载所有测试数据预测时,混淆矩阵准确率符合预期(约90%),但如果无法一次性获取全部测试数据,找不到小批量生成器失效的原因,需要帮助。

我基于COVID-19 Radiography Dataset构建二分类图像分类器,区分COVID感染肺部图像和正常肺部图像。为了利用数据集附带的肺部掩码,我写了一个自定义数据生成器,以批量方式将肺部图像和掩码逐元素相乘处理。

代码运行看似正常,但出现了矛盾结果:

  • 模型evaluate输出的测试准确率约91%
  • 用scikit-learn生成的混淆矩阵、手动计算的准确率仅约62%
  • 手动抽查100个预测结果,错误数约9个,和evaluate的准确率一致

排查过生成器的shuffle问题,即使全程关闭shuffle,结果依然相同(当前设置训练时打乱,测试时不打乱)。


自定义数据生成器代码

def custom_generator(image_paths, mask_paths, labels, batch_size=32, shuffle=True): 
    num_samples = len(image_paths) 
    while True:  # Keras生成器需要无限循环
        indices = np.arange(num_samples) 
        if shuffle: np.random.shuffle(indices)    
    
        for start in range(0, num_samples, batch_size):
            end = min(start + batch_size, num_samples)  # 正确处理最后一批的索引
            batch_indices = indices[start:end]  # 选择当前批次的索引
            
            # 诊断用打印(已注释)
            """
            print(f"Generating batch: {start // batch_size + 1}")
            print(f"Start index: {start}, End index: {end}")
            print(f"Number of indices in this batch: {len(batch_indices)}")
            """
            batch_images = []
            batch_labels = []
            
            for idx in batch_indices:
                img_path = image_paths[int(idx)]
                mask_path = mask_paths[int(idx)]
                
                # 加载并处理图像和掩码
                image = Image.open(img_path).convert("L")
                mask = Image.open(mask_path).convert("L")
                mask = mask.resize((299, 299), Image.LANCZOS)
                
                image_array = np.array(image) / 255.0
                mask_array = np.array(mask) / 255.0
                
                multiplied_image = image_array * mask_array
                
                # 扩展维度添加通道
                batch_images.append(np.expand_dims(multiplied_image, axis=-1))
                batch_labels.append(labels[int(idx)])
            
            yield np.array(batch_images), np.array(batch_labels)  

模型评估与混淆矩阵代码

# 初始化测试生成器
test_generator = custom_generator(test_image_paths, test_mask_paths, test_labels, batch_size=1, shuffle=False)

test_steps = len(test_image_paths)  
test_loss, test_accuracy = model.evaluate(test_generator, steps=test_steps)

print(f"Test Accuracy: {test_accuracy * 100:.2f}%")
print(f"Test Loss: {test_loss:.4f}")

# 生成预测结果
test_predictions = model.predict(test_generator, steps=test_steps)
predicted_classes = (test_predictions > 0.5).astype("int32")

# 重新初始化生成器(这里有问题)
test_generator = custom_generator(test_image_paths, test_mask_paths, test_labels, batch_size=1)

# 计算混淆矩阵和准确率
from sklearn.metrics import confusion_matrix, accuracy_score

cm = confusion_matrix(test_labels, predicted_classes)

print("Confusion Matrix:")
print(cm)

cm_accuracy = np.trace(cm) / np.sum(cm)
print(f"Accuracy from Confusion Matrix: {cm_accuracy * 100:.2f}%")

manual_accuracy = accuracy_score(test_labels, predicted_classes)
print(f"Manually Calculated Accuracy: {manual_accuracy * 100:.2f}%")

问题分析与解决方案

核心问题:生成器的状态不一致

Keras的生成器是有状态的迭代器,model.evaluate会耗尽生成器的迭代状态,之后直接调用model.predict时,生成器会从上次结束的位置继续迭代,而不是从头开始。这就导致predicted_classes的顺序和test_labels完全不匹配,混淆矩阵自然错误。

另外,你后续还重新初始化了一个带shuffle=True的生成器,但这步完全多余,反而加剧了混乱。

修复步骤

  1. 避免重复使用同一个生成器实例:每次调用evaluate或predict前,都要重新初始化生成器,确保从数据集开头开始迭代
  2. 统一生成器的shuffle设置:测试阶段必须关闭shuffle,保证预测结果和标签的顺序严格对应

修改后的评估代码:

# 评估模型:重新初始化生成器
test_generator_eval = custom_generator(test_image_paths, test_mask_paths, test_labels, batch_size=1, shuffle=False)
test_steps = len(test_image_paths)  
test_loss, test_accuracy = model.evaluate(test_generator_eval, steps=test_steps)

print(f"Test Accuracy: {test_accuracy * 100:.2f}%")
print(f"Test Loss: {test_loss:.4f}")

# 生成预测:再次初始化全新的生成器
test_generator_pred = custom_generator(test_image_paths, test_mask_paths, test_labels, batch_size=1, shuffle=False)
test_predictions = model.predict(test_generator_pred, steps=test_steps)
predicted_classes = (test_predictions > 0.5).astype("int32")

# 计算混淆矩阵和准确率
from sklearn.metrics import confusion_matrix, accuracy_score

cm = confusion_matrix(test_labels, predicted_classes)

print("Confusion Matrix:")
print(cm)

cm_accuracy = np.trace(cm) / np.sum(cm)
print(f"Accuracy from Confusion Matrix: {cm_accuracy * 100:.2f}%")

manual_accuracy = accuracy_score(test_labels, predicted_classes)
print(f"Manually Calculated Accuracy: {manual_accuracy * 100:.2f}%")

针对无法一次性加载数据的场景优化

如果测试集太大无法一次性加载,可以:

  • 在生成器迭代时同步收集真实标签,确保预测和标签的顺序完全对应
  • 或者将生成器的输出缓存到磁盘(如分批次保存为.npy文件),再批量处理预测结果

示例:同步收集标签的方式

test_generator = custom_generator(test_image_paths, test_mask_paths, test_labels, batch_size=1, shuffle=False)

# 同时收集预测结果和真实标签
predicted_classes = []
true_labels = []
for _ in range(test_steps):
    imgs, lbls = next(test_generator)
    pred = model.predict(imgs, verbose=0)
    predicted_classes.append((pred > 0.5).astype("int32")[0][0])
    true_labels.append(lbls[0][0])

# 现在计算混淆矩阵
cm = confusion_matrix(true_labels, predicted_classes)
# ... 后续计算准确率

这样即使生成器有状态问题,也能保证预测和标签一一对应,不会出现顺序错位。

内容的提问来源于stack exchange,提问作者Dawson Kimyon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 00:39:51