You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Inria航空影像数据集的Unet模型训练报错:ValueError——`logits`与`labels`形状不匹配

解决Unet实例分割中Logits与Labels形状不匹配的问题

嘿,我一眼就看到你的问题出在哪了——从报错信息里的((None, 1024, 1024, 1) vs (None, 1024, 1024, 3))就能明确:你的模型输出是单通道的二分类结果,但训练标签Y_train被处理成了3通道,两者形状完全不匹配,这才导致了损失计算时的错误。

问题根源

你写的resize2D函数对RGB图像和掩码图像用了同样的处理逻辑:用cv2.IMREAD_COLOR读取所有图像。但你的掩码是单通道灰度图(像素0和1),用彩色模式读取后,OpenCV会自动把它转换成3通道的RGB图(三个通道的像素值完全一样),最终Y_train的形状就变成了(样本数, 1024, 1024, 3);而你用的Unet模型在二分类任务下,默认输出是单通道(样本数, 1024, 1024, 1),两者形状对不上,自然报错。

修复方案

我们需要分别处理RGB训练图和掩码图,确保掩码是单通道的4D张量(符合TensorFlow的输入要求),具体修改如下:

1. 拆分图像与掩码的处理函数

把原来的resize2D拆成两个函数,分别处理RGB图像和掩码:

import tensorflow as tf
from tensorflow import keras
import segmentation_models as sm
import glob
import cv2
import os
import numpy as np
from matplotlib import pyplot as plt

# Set directory paths to X and Y training data
Xtrain_path = 'DIRECTORY'
Ytrain_path = 'DIRECTORY'

# Set backbone model and preprocessing instance through segmentation_models
BACKBONE = 'resnet34'
preprocessed_input = sm.get_preprocessing(BACKBONE)

# 处理RGB训练图像的resize函数
def resize_rgb(directory, target_size=(1024, 1024)):
    img_array = []
    x, y = target_size
    for filename in next(os.walk(directory))[2]:
        img_path = os.path.join(directory, filename)
        # 读取彩色图像并转成RGB格式
        img = cv2.imread(img_path, cv2.IMREAD_COLOR)
        img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)
        img = cv2.resize(img, (y, x))
        img_array.append(img)
    return np.array(img_array)

# 处理掩码图像的resize函数
def resize_mask(directory, target_size=(1024, 1024)):
    img_array = []
    x, y = target_size
    for filename in next(os.walk(directory))[2]:
        img_path = os.path.join(directory, filename)
        # 以灰度模式读取掩码(单通道)
        img = cv2.imread(img_path, cv2.IMREAD_GRAYSCALE)
        img = cv2.resize(img, (y, x))
        # 确保掩码像素值是0-1(如果原始掩码是0和255的话)
        img = img / 255.0
        # 扩展通道维度,变成4D张量(符合模型输入要求)
        img = np.expand_dims(img, axis=-1)
        img_array.append(img)
    return np.array(img_array)

# 加载并处理数据
X_train = resize_rgb(directory=Xtrain_path)
Y_train = resize_mask(directory=Ytrain_path)

# Preprocess X_train data
X_train = preprocessed_input(X_train)

# 设置模型框架,使用imagenet预训练权重
model = sm.Unet(BACKBONE, encoder_weights="imagenet")

# 编译模型,二分类任务使用binary_crossentropy损失
model.compile(optimizer='adam',loss='binary_crossentropy',metrics=['mse'])
print(model.summary())

# 启动训练
run1 = model.fit(X_train,Y_train,validation_split=0.1,batch_size=10,epochs=20,verbose=1)

2. 关键修改点说明

  • 掩码读取模式:用cv2.IMREAD_GRAYSCALE读取掩码,确保得到单通道数组
  • 归一化掩码:如果你的原始掩码是8位灰度图(像素0和255),除以255.0把值缩到0-1区间,和binary_crossentropy的要求匹配
  • 扩展通道维度:用np.expand_dims把3D的掩码数组变成4D张量,形状从(N,1024,1024)变成(N,1024,1024,1),和模型输出形状一致

这样修改后,你的模型就能正常计算损失并开始训练了。另外,你选择的ResNet34+Unet+Binary Crossentropy的组合对于建筑足迹分割是完全合适的,没问题。

内容的提问来源于stack exchange,提问作者mdickson4

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 20:14:04