You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

复刻CNN基准模型无法达到0.5得分的问题排查求助

复刻CNN基准模型无法达到0.5得分的问题排查求助

我最近在做一个图像分割项目,有个公开的基准模型能达到0.5的IoU得分,但我严格照着它的参数和流程复现的时候,模型最多只能拿到0.3的IoU,试了各种调整都没法突破这个分数,实在摸不着头绪,想请各位帮忙排查下问题出在哪。

先跟大家说下数据和基准的基本情况:

  • 掩码是36×36的二值图像,最终会保存为CSV格式
  • 基准模型是一个简单的CNN架构,专门在从标准数据集提取的36×36图像补丁上训练,模型末尾加了sigmoid激活函数

下面是我严格遵循的预处理步骤:

  • 用0填充所有缺失值
  • 从训练集中系统性移除包含异常值的补丁
  • 对每个补丁单独使用RobustScaler做归一化,保证缩放一致性

训练的超参数我也完全对齐了:

  • CNN包含5层卷积层,没有池化操作
  • 批大小设为128,优化训练时的计算效率
  • 学习率设置为0.001,引导模型高效收敛
  • 训练30个epoch,充分捕捉数据特征的变化
  • 使用二元交叉熵(Binary Cross Entropy)作为损失函数,平衡Dice系数和BCE的优化效果
  • 加入了水平翻转和水平滚动的数据增强,提升模型的泛化能力
  • 优化器选用Adam

我把完整的代码贴在下面,麻烦各位帮忙看看有没有哪里遗漏或者出错的地方:

import os
import numpy as np
from tqdm import tqdm
from sklearn.preprocessing import RobustScaler
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Conv2D, Flatten, Dense, Activation
from tensorflow.keras.optimizers import Adam
from tensorflow.keras.losses import BinaryCrossentropy
import pandas as pd

# Collect values for percentile calculation
all_values = []
for file_name in tqdm(sorted(os.listdir(x_train_dir))):
    if file_name.endswith('.npy'):
        img = np.load(os.path.join(x_train_dir, file_name))
        img[np.isnan(img)] = 0  # Replace NaN with 0
        all_values.extend(img.flatten())

# Calculate percentiles for filtering
lower_bound = np.percentile(all_values, 1)
upper_bound = np.percentile(all_values, 99)

# Load and filter training data
x_train = []
y_train_processed = []

for file_name in tqdm(sorted(os.listdir(x_train_dir))):
    if file_name.endswith('.npy'):
        img = np.load(os.path.join(x_train_dir, file_name))
        img[np.isnan(img)] = 0  # Replace NaN with 0
        
        # Filter outliers
        if np.all((img >= lower_bound) & (img <= upper_bound)):
            scaler = RobustScaler()
            img = scaler.fit_transform(img)  # Normalization
            
            # Apply augmentations
            augmented_images = [img]  # Original
            augmented_images.append(np.fliplr(img))  # Horizontal flip
            augmented_images.append(np.roll(img, shift=10, axis=1))  # Horizontal roll (10 pixels)

            # Add each augmentation to the training set
            for augmented_img in augmented_images:
                x_train.append(augmented_img)
                
                # Load corresponding mask (same for all augmentations)
                patch_name = file_name.replace('.npy', '')
                y_train_processed.append(y_train.loc[patch_name].values.reshape((36, 36)))

# Convert to numpy array
x_train = np.array(x_train).reshape(-1, 36, 36, 1)
y_train_processed = np.array(y_train_processed).reshape(-1, 36, 36, 1)

print("Loading, filtering, and augmenting training data completed.")

# Load and preprocess test data
x_test = []
test_file_names = []

# Apply the same scaler as for the training data
for file_name in tqdm(sorted(os.listdir(x_test_dir))):
    if file_name.endswith('.npy'):
        # Load image
        img = np.load(os.path.join(x_test_dir, file_name))
        img[np.isnan(img)] = 0  # Replace NaN with 0
        
        # Clipping outliers based on the bounds calculated from training data
        img = np.clip(img, lower_bound, upper_bound)
        
        # Apply the scaler (pre-adjusted on training data)
        img = scaler.transform(img)
        
        # Add to dataset
        x_test.append(img)
        test_file_names.append(file_name.replace('.npy', ''))

# Convert to numpy array
x_test = np.array(x_test).reshape(-1, 36, 36, 1)

print("Test data preprocessing completed.")

# Define the CNN model
model = Sequential([
    Conv2D(32, (3, 3), activation='relu', input_shape=(36, 36, 1)),
    Conv2D(64, (3, 3), activation='relu'),
    Conv2D(128, (3, 3), activation='relu'),
    Conv2D(64, (3, 3), activation='relu'),
    Conv2D(32, (3, 3), activation='relu'),
    Flatten(),
    Dense(36 * 36, activation='sigmoid'),
    Activation('sigmoid')
])

# Compile the model
model.compile(optimizer=Adam(learning_rate=0.001),
              loss=BinaryCrossentropy(),
              metrics=['accuracy'])

# Train the model (guaranteed 30 epochs)
history = model.fit(
    x_train,
    y_train_processed.reshape(-1, 36 * 36),
    batch_size=128,
    epochs=30,
    verbose=1
)

# Function to calculate IoU
def calculate_iou(y_true, y_pred):
    """Calculate IoU between true mask and predicted mask."""
    intersection = np.logical_and(y_true, y_pred).sum()  # Intersection
    union = np.logical_or(y_true, y_pred).sum()          # Union
    iou = intersection / union if union != 0 else 1.0    # Handle cases where union = 0
    return iou

# Evaluation on training data
predictions_train = model.predict(x_train)
ious = []  # List to store IoU
for i, pred in enumerate(predictions_train):
    y_pred = (pred > 0.5).astype(int).reshape(36, 36)  # Binarize the prediction
    y_true = y_train_processed[i].reshape(36, 36)     # Ground truth
    iou = calculate_iou(y_true, y_pred)
    ious.append(iou)

# Mean IoU result
mean_iou = np.mean(ious)
print(f"Mean IoU score on training data: {mean_iou:.4f}")

# Predictions on test data
test_predictions = model.predict(x_test)

# Create the submission file
submission = []
for i, patch_name in enumerate(test_file_names):  # Use test file names
    y_pred = (test_predictions[i] > 0.5).astype(int).flatten()
    submission.append([patch_name] + y_pred.tolist())

# Save in the expected format
submission_df = pd.DataFrame(submission, columns=['Unnamed: 0'] + [str(i) for i in range(1296)])
submission_df.to_csv('submissionV2.csv', index=False)

print("Submission file created successfully.")

另外补充下,我在代码里已经注意对齐了所有基准提到的步骤,但训练出来的IoU就是卡在0.3左右,训练集上的IoU也不高,感觉模型根本没学到东西?或者是不是我哪里预处理或者模型定义错了?

备注:内容来源于stack exchange,提问作者sandrowin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.16 02:55:31