You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

语义分割输出全黑无掩码,准确率固定但损失下降问题求助

语义分割训练异常:全黑输出、准确率固定但损失下降

我用以下代码在自定义数据集(40张图片)和COCO格式标注(仅人脸1类)上执行语义分割时,出现了以下问题:

  • 模型输出全黑图像,无有效掩码,所有像素值一致
  • 训练过程中每轮准确率固定为67.87%,但损失值却逐轮下降

训练代码

import cv2
import os
import json
import numpy as np
import tensorflow as tf
from tensorflow import keras
from tensorflow.keras.layers import Input, Dense, Conv2D, MaxPooling2D, UpSampling2D
from tensorflow.keras.models import Model
from pycocotools.coco import COCO
from tensorflow.keras.applications import VGG16
from sklearn.utils import compute_sample_weight

folder_path = "D:\ImageClassification\face_semantic_segmentation\dataset"
filenames = os.listdir(folder_path)
images = []
for filename in filenames:
    img = cv2.imread(os.path.join(folder_path, filename))
    img = cv2.resize(img, (256, 256))
    img = np.array(img)
    images.append(img)
x_train = np.array(images)

with open("D:\ImageClassification\face_semantic_segmentation\annotations.json") as f:
    coco = json.load(f)
annotations = coco['annotations']
masks = {}
for annotation in annotations:
    image_id = annotation['image_id']
    if image_id not in masks:
        masks[image_id] = []
    masks[image_id].append(annotation['segmentation'])

resized_masks = []
for image_id, mask in masks.items():
    mask_img = np.zeros((720, 1280), dtype=np.uint8)
    for segmentation in mask:
        poly = np.array(segmentation).reshape((-1, 1, 2)).astype(np.int32)
        cv2.fillPoly(mask_img, [poly], 1)
    mask_img = cv2.resize(mask_img, (256, 256))
    mask_img = np.stack([mask_img] * 3, axis=-1)
    resized_masks.append(mask_img)
y_train = np.array(resized_masks)

import matplotlib.pyplot as plt
mask = y_train[4]
mask = np.sum(mask, axis=-1)
plt.imshow(mask)
plt.show()

from segmentation_models import Unet
from segmentation_models import get_preprocessing
from segmentation_models.losses import bce_jaccard_loss
from segmentation_models.metrics import iou_score

BACKBONE = 'resnet50'
preprocess_input = get_preprocessing(BACKBONE)

model = Unet(BACKBONE,classes=2,input_shape=(256,256, 3), encoder_weights='imagenet',activation='sigmoid')
x_train = preprocess_input(x_train)
x = model.layers[-1].output
x = Conv2D(3, (1, 1), activation='sigmoid')(x)
model = Model(inputs=model.input, outputs=x)
model.compile(optimizer='Adam', loss='binary_crossentropy', metrics=['binary_accuracy'])
model.fit(x=x_train,y=y_train,batch_size=32,epochs=50)

问题根源分析

  • 图像与标注不匹配:通过os.listdir读取图像的顺序,和COCO标注中image_id对应的图像顺序完全无关,导致训练时图像和掩码完全错位,模型学到的是无意义的映射。
  • 掩码处理错误:将单通道掩码堆叠为3通道,违背单类分割的任务逻辑;同时固定用720x1280初始化掩码,若原始图像分辨率不同,会导致多边形绘制错位,生成无效掩码。
  • 模型输出层错误:原Unet已设置classes=2,却额外添加Conv2D(3)输出3通道,和任务目标不匹配,混淆了损失计算逻辑。
  • 数据不平衡导致指标异常:人脸在图像中占比通常远小于背景,模型直接预测背景就能稳定拿到67.87%的准确率,损失下降是因为模型在优化这个错误的预测方向。

修复方案

1. 修正图像与标注的匹配逻辑

利用COCO标注中images字段的映射关系,确保图像和掩码一一对应:

# 建立image_id到文件名的映射
image_id_to_filename = {img['id']: img['file_name'] for img in coco['images']}

# 按标注的image_id读取对应图像
images = []
valid_image_ids = []
for image_id in masks.keys():
    if image_id in image_id_to_filename:
        filename = image_id_to_filename[image_id]
        img_path = os.path.join(folder_path, filename)
        img = cv2.imread(img_path)
        if img is not None:
            img = cv2.resize(img, (256, 256))
            images.append(img)
            valid_image_ids.append(image_id)

# 生成对应顺序的掩码
resized_masks = []
for image_id in valid_image_ids:
    # 获取原始图像的真实分辨率
    orig_img = next(img for img in coco['images'] if img['id'] == image_id)
    orig_h, orig_w = orig_img['height'], orig_img['width']
    mask_img = np.zeros((orig_h, orig_w), dtype=np.uint8)
    
    for segmentation in masks[image_id]:
        poly = np.array(segmentation).reshape((-1, 1, 2)).astype(np.int32)
        cv2.fillPoly(mask_img, [poly], 1)
    
    mask_img = cv2.resize(mask_img, (256, 256))
    # 单类分割保留单通道
    resized_masks.append(mask_img[..., np.newaxis])

x_train = np.array(images)
y_train = np.array(resized_masks)

2. 修正模型输出与配置

单类语义分割应输出1通道(二分类),使用分割专用的损失和指标:

BACKBONE = 'resnet50'
preprocess_input = get_preprocessing(BACKBONE)

# 配置正确的Unet模型
model = Unet(
    BACKBONE,
    classes=1,  # 单类二分类,输出1通道
    input_shape=(256, 256, 3),
    encoder_weights='imagenet',
    activation='sigmoid'
)

# 使用更适合分割的损失函数和指标
model.compile(
    optimizer='Adam',
    loss=bce_jaccard_loss,  # 结合交叉熵和Jaccard损失,比纯交叉熵更适合分割
    metrics=[iou_score, 'binary_accuracy']
)

3. 处理数据不平衡

通过样本权重提升前景(人脸)的训练权重,避免模型偏向背景:

# 生成像素级样本权重,前景权重设为5,背景为1
sample_weights = []
for mask in y_train:
    weight = np.where(mask == 1, 5, 1)
    sample_weights.append(weight)
sample_weights = np.array(sample_weights)

# 调整batch_size,40张图用32太大,改为8更合理
model.fit(
    x=preprocess_input(x_train),
    y=y_train,
    batch_size=8,
    epochs=50,
    sample_weight=sample_weights
)

4. 验证掩码正确性

训练前务必验证图像和掩码的对应关系,避免无效数据:

# 随机选取一组数据验证
idx = np.random.randint(0, len(x_train))
plt.figure(figsize=(10, 5))
plt.subplot(121)
plt.imshow(cv2.cvtColor(x_train[idx], cv2.COLOR_BGR2RGB))
plt.title('原始图像')
plt.subplot(122)
plt.imshow(y_train[idx][..., 0], cmap='gray')
plt.title('对应掩码')
plt.show()

内容的提问来源于stack exchange,提问作者Ahmad Hamayel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 15:45:23