You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否手动创建混淆矩阵的y_true数组?如何从CustomDataGenerator获取y_true

手动创建验证集标签数组的可行性及优化方案

完全可以像你这样手动构建y_true数组,只要你能保证验证集样本的顺序和你定义的标签顺序完全对应——而你已经设置了shuffle=False,同时明确了每个标签的样本数量,这样计算出来的混淆矩阵结果是准确可信的。

不过可以优化一下手动构建数组的写法,避免逐个输入数字的繁琐和出错概率:

# 根据各标签样本数量批量生成y_true
y_true = np.concatenate([
    np.full(12, 0),   # 12个label0
    np.full(6, 1),    # 6个label1
    np.full(7, 2),    # 7个label2
    np.full(18, 3),   # 18个label3
    np.full(4, 4),    # 4个label4
    np.full(10, 5),   # 10个label5
    np.full(4, 6)     # 4个label6
])

如果你想直接从CustomDataGenerator中获取y_true,可以根据生成器的实现逻辑来收集标签:
假设你的生成器在__getitem__方法中返回(图像数据, 标签),可以通过遍历生成器来收集所有标签:

y_true = []
for batch_x, batch_y in val_datagen:
    # 如果标签是one-hot编码格式,先转成类别索引
    if batch_y.ndim > 1:
        batch_labels = np.argmax(batch_y, axis=1)
    else:
        batch_labels = batch_y
    y_true.extend(batch_labels)
y_true = np.array(y_true)

这种方式不需要手动统计样本数量,更灵活,适合后续验证集样本数量有变动的场景。

内容的提问来源于stack exchange,提问作者Syuuuu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 18:33:24