You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

土地覆盖分类:13通道影像适配预训练模型的维度缩减方案咨询

土地覆盖分类任务中13波段TIFF适配预训练模型的解决方案

问题背景

做土地覆盖分类任务,测试预训练模型时,用EuroSAT数据集的RGB(3通道)JPG图像完全正常,但处理13波段的TIFF图像卡壳了:预训练模型要求输入必须是3通道,13通道数据直接喂进去根本跑不起来。试过用PCA降维把(64,64,13)转成(64,64,3),但tf.pca只支持2D张量输入,这条路走不通。

现有代码

数据集划分代码

ds_train = tf.data.Dataset.from_tensor_slices((df_train_strat['images_path'], df_train_strat["encoded_label"]))
ds_train = ds_train.map(lambda x, y : tf.py_function(parse_image, [x, y], [tf.float32, tf.int64]))
ds_train = ds_train.map(_fixup_shape)
ds_train = ds_train.batch(BATCH_SIZE)

ds_test = tf.data.Dataset.from_tensor_slices((df_test_strat['images_path'], df_test_strat["encoded_label"]))
ds_test = ds_test.map(lambda x, y : tf.py_function(parse_image, [x, y], [tf.float32, tf.int64]))
ds_test = ds_test.map(_fixup_shape)
ds_test = ds_test.batch(BATCH_SIZE)

图像解析函数

def parse_image(img_path: str, label: str):
    # Cast the Tensor to numpy and decode the string
    img_path = img_path.numpy().decode('utf-8')

    with rasterio.open(img_path) as src:
        img = src.read()
        # Channels last
        img = np.moveaxis(img, 0, 2)
        # Images normalization
        array_min, array_max = img.min(), img.max()
        img = (img - array_min)/(array_max - array_min)

    return img, label

解析后张量形状为(64,64,64,13)(批量大小为64)。

预训练模型调用与模型构建

base_model = EfficientNetB0(weights='imagenet', include_top=False, input_shape=(64, 64, 3))

这里改input_shape就报错,没法直接调整。

model = models.Sequential()
model.add(base_model)
model.add(layers.GlobalMaxPooling2D())

model.add(layers.Dense(512, activation='relu', kernel_initializer="he_normal"))
model.add(layers.BatchNormalization())
model.add(layers.Dropout(0.5))
model.add(layers.Dense(128, activation='relu', kernel_initializer="he_normal"))
model.add(layers.BatchNormalization())
model.add(layers.Dropout(0.5))

model.add(layers.Dense(10, activation = 'softmax', kernel_initializer="glorot_normal"))

model.summary()

训练代码

early_stop = callbacks.EarlyStopping(monitor = 'val_loss', mode = 'min',
            patience = 3, restore_best_weights = True, verbose = 1)

reduce_lr = callbacks.ReduceLROnPlateau(monitor = 'val_loss', mode = 'min',
            patience = 2, factor = 0.5, min_lr = 1e-06, verbose = 1)

    
model.compile(optimizer=adam, loss='categorical_crossentropy',
              metrics=['accuracy'])

history = model.fit(ds_train, validation_data=ds_test, epochs=20, batch_size = 64,
                    callbacks=[reduce_lr, early_stop])

训练时核心报错:输入张量通道数(13)与模型预期的3通道不匹配。

可行解决方案

方案1:加1x1卷积层适配通道数

在预训练模型前面加一个1x1卷积层,把13通道直接映射成3通道,既能复用预训练模型的权重,又不用改数据集。修改模型构建代码如下:

# 定义输入层,适配13通道
input_layer = layers.Input(shape=(64, 64, 13))
# 1x1卷积做通道转换,参数可训练
channel_adapt = layers.Conv2D(3, (1,1), activation='relu', kernel_initializer="he_normal")(input_layer)

base_model = EfficientNetB0(weights='imagenet', include_top=False, input_shape=(64, 64, 3))
# 先冻结预训练层,训练后期可根据效果解冻微调
base_model.trainable = False

x = base_model(channel_adapt)
x = layers.GlobalMaxPooling2D()(x)

x = layers.Dense(512, activation='relu', kernel_initializer="he_normal")(x)
x = layers.BatchNormalization()(x)
x = layers.Dropout(0.5)(x)
x = layers.Dense(128, activation='relu', kernel_initializer="he_normal")(x)
x = layers.BatchNormalization()(x)
x = layers.Dropout(0.5)(x)

output_layer = layers.Dense(10, activation = 'softmax', kernel_initializer="glorot_normal")(x)

model = models.Model(inputs=input_layer, outputs=output_layer)
model.summary()

方案2:手动实现多通道图像的PCA降维

既然tf.pca不好用,就用sklearn的PCA手动处理。可以选择单张图像独立降维,或者更严谨的——先计算整个训练集的PCA参数再统一应用,避免数据泄露。

方式A:单张图像独立PCA(简单但效果一般)

修改parse_image函数:

from sklearn.decomposition import PCA

def parse_image(img_path: str, label: str):
    img_path = img_path.numpy().decode('utf-8')

    with rasterio.open(img_path) as src:
        img = src.read()
        img = np.moveaxis(img, 0, 2)
        # 归一化
        array_min, array_max = img.min(), img.max()
        img = (img - array_min)/(array_max - array_min)
        
        # 把(64,64,13)转成(64*64,13)做PCA
        h, w, c = img.shape
        img_reshaped = img.reshape(-1, c)
        pca = PCA(n_components=3)
        img_pca = pca.fit_transform(img_reshaped)
        # 还原回图像形状
        img_pca = img_pca.reshape(h, w, 3)
        # 再次归一化到[0,1]保证数据稳定
        img_pca = (img_pca - img_pca.min())/(img_pca.max() - img_pca.min())

    return img_pca.astype(np.float32), label

方式B:全局PCA(更合理,避免数据泄露)

先遍历所有训练图像计算全局PCA参数,再应用到训练和测试集:

from sklearn.decomposition import PCA

# 先收集训练集所有像素用于计算PCA
def collect_all_pixels(df):
    all_pixels = []
    for img_path in df['images_path']:
        with rasterio.open(img_path) as src:
            img = src.read()
            img = np.moveaxis(img, 0, 2)
            array_min, array_max = img.min(), img.max()
            img = (img - array_min)/(array_max - array_min)
            all_pixels.append(img.reshape(-1, 13))
    return np.concatenate(all_pixels, axis=0)

# 计算训练集的PCA
train_pixels = collect_all_pixels(df_train_strat)
pca = PCA(n_components=3)
pca.fit(train_pixels)

# 修改解析函数,用预训练好的PCA
def parse_image_pca(img_path: str, label: str):
    img_path = img_path.numpy().decode('utf-8')

    with rasterio.open(img_path) as src:
        img = src.read()
        img = np.moveaxis(img, 0, 2)
        array_min, array_max = img.min(), img.max()
        img = (img - array_min)/(array_max - array_min)
        
        h, w, c = img.shape
        img_reshaped = img.reshape(-1, c)
        img_pca = pca.transform(img_reshaped)
        img_pca = img_pca.reshape(h, w, 3)
        img_pca = (img_pca - img_pca.min())/(img_pca.max() - img_pca.min())

    return img_pca.astype(np.float32), label

之后数据集映射时换成parse_image_pca即可。

方案3:换用支持多通道的预训练模型

找针对遥感图像预训练的模型(比如在SENet、ResNet基础上用遥感数据集预训练的版本),或者直接从头训练模型。但这种方法会丢失ImageNet预训练权重的优势,训练时间和成本更高。

内容的提问来源于stack exchange,提问作者Brocoli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 19:05:18