土地覆盖分类:13通道影像适配预训练模型的维度缩减方案咨询
土地覆盖分类任务中13波段TIFF适配预训练模型的解决方案
问题背景
做土地覆盖分类任务,测试预训练模型时,用EuroSAT数据集的RGB(3通道)JPG图像完全正常,但处理13波段的TIFF图像卡壳了:预训练模型要求输入必须是3通道,13通道数据直接喂进去根本跑不起来。试过用PCA降维把(64,64,13)转成(64,64,3),但tf.pca只支持2D张量输入,这条路走不通。
现有代码
数据集划分代码
ds_train = tf.data.Dataset.from_tensor_slices((df_train_strat['images_path'], df_train_strat["encoded_label"])) ds_train = ds_train.map(lambda x, y : tf.py_function(parse_image, [x, y], [tf.float32, tf.int64])) ds_train = ds_train.map(_fixup_shape) ds_train = ds_train.batch(BATCH_SIZE) ds_test = tf.data.Dataset.from_tensor_slices((df_test_strat['images_path'], df_test_strat["encoded_label"])) ds_test = ds_test.map(lambda x, y : tf.py_function(parse_image, [x, y], [tf.float32, tf.int64])) ds_test = ds_test.map(_fixup_shape) ds_test = ds_test.batch(BATCH_SIZE)
图像解析函数
def parse_image(img_path: str, label: str): # Cast the Tensor to numpy and decode the string img_path = img_path.numpy().decode('utf-8') with rasterio.open(img_path) as src: img = src.read() # Channels last img = np.moveaxis(img, 0, 2) # Images normalization array_min, array_max = img.min(), img.max() img = (img - array_min)/(array_max - array_min) return img, label
解析后张量形状为(64,64,64,13)(批量大小为64)。
预训练模型调用与模型构建
base_model = EfficientNetB0(weights='imagenet', include_top=False, input_shape=(64, 64, 3))
这里改input_shape就报错,没法直接调整。
model = models.Sequential() model.add(base_model) model.add(layers.GlobalMaxPooling2D()) model.add(layers.Dense(512, activation='relu', kernel_initializer="he_normal")) model.add(layers.BatchNormalization()) model.add(layers.Dropout(0.5)) model.add(layers.Dense(128, activation='relu', kernel_initializer="he_normal")) model.add(layers.BatchNormalization()) model.add(layers.Dropout(0.5)) model.add(layers.Dense(10, activation = 'softmax', kernel_initializer="glorot_normal")) model.summary()
训练代码
early_stop = callbacks.EarlyStopping(monitor = 'val_loss', mode = 'min', patience = 3, restore_best_weights = True, verbose = 1) reduce_lr = callbacks.ReduceLROnPlateau(monitor = 'val_loss', mode = 'min', patience = 2, factor = 0.5, min_lr = 1e-06, verbose = 1) model.compile(optimizer=adam, loss='categorical_crossentropy', metrics=['accuracy']) history = model.fit(ds_train, validation_data=ds_test, epochs=20, batch_size = 64, callbacks=[reduce_lr, early_stop])
训练时核心报错:输入张量通道数(13)与模型预期的3通道不匹配。
可行解决方案
方案1:加1x1卷积层适配通道数
在预训练模型前面加一个1x1卷积层,把13通道直接映射成3通道,既能复用预训练模型的权重,又不用改数据集。修改模型构建代码如下:
# 定义输入层,适配13通道 input_layer = layers.Input(shape=(64, 64, 13)) # 1x1卷积做通道转换,参数可训练 channel_adapt = layers.Conv2D(3, (1,1), activation='relu', kernel_initializer="he_normal")(input_layer) base_model = EfficientNetB0(weights='imagenet', include_top=False, input_shape=(64, 64, 3)) # 先冻结预训练层,训练后期可根据效果解冻微调 base_model.trainable = False x = base_model(channel_adapt) x = layers.GlobalMaxPooling2D()(x) x = layers.Dense(512, activation='relu', kernel_initializer="he_normal")(x) x = layers.BatchNormalization()(x) x = layers.Dropout(0.5)(x) x = layers.Dense(128, activation='relu', kernel_initializer="he_normal")(x) x = layers.BatchNormalization()(x) x = layers.Dropout(0.5)(x) output_layer = layers.Dense(10, activation = 'softmax', kernel_initializer="glorot_normal")(x) model = models.Model(inputs=input_layer, outputs=output_layer) model.summary()
方案2:手动实现多通道图像的PCA降维
既然tf.pca不好用,就用sklearn的PCA手动处理。可以选择单张图像独立降维,或者更严谨的——先计算整个训练集的PCA参数再统一应用,避免数据泄露。
方式A:单张图像独立PCA(简单但效果一般)
修改parse_image函数:
from sklearn.decomposition import PCA def parse_image(img_path: str, label: str): img_path = img_path.numpy().decode('utf-8') with rasterio.open(img_path) as src: img = src.read() img = np.moveaxis(img, 0, 2) # 归一化 array_min, array_max = img.min(), img.max() img = (img - array_min)/(array_max - array_min) # 把(64,64,13)转成(64*64,13)做PCA h, w, c = img.shape img_reshaped = img.reshape(-1, c) pca = PCA(n_components=3) img_pca = pca.fit_transform(img_reshaped) # 还原回图像形状 img_pca = img_pca.reshape(h, w, 3) # 再次归一化到[0,1]保证数据稳定 img_pca = (img_pca - img_pca.min())/(img_pca.max() - img_pca.min()) return img_pca.astype(np.float32), label
方式B:全局PCA(更合理,避免数据泄露)
先遍历所有训练图像计算全局PCA参数,再应用到训练和测试集:
from sklearn.decomposition import PCA # 先收集训练集所有像素用于计算PCA def collect_all_pixels(df): all_pixels = [] for img_path in df['images_path']: with rasterio.open(img_path) as src: img = src.read() img = np.moveaxis(img, 0, 2) array_min, array_max = img.min(), img.max() img = (img - array_min)/(array_max - array_min) all_pixels.append(img.reshape(-1, 13)) return np.concatenate(all_pixels, axis=0) # 计算训练集的PCA train_pixels = collect_all_pixels(df_train_strat) pca = PCA(n_components=3) pca.fit(train_pixels) # 修改解析函数,用预训练好的PCA def parse_image_pca(img_path: str, label: str): img_path = img_path.numpy().decode('utf-8') with rasterio.open(img_path) as src: img = src.read() img = np.moveaxis(img, 0, 2) array_min, array_max = img.min(), img.max() img = (img - array_min)/(array_max - array_min) h, w, c = img.shape img_reshaped = img.reshape(-1, c) img_pca = pca.transform(img_reshaped) img_pca = img_pca.reshape(h, w, 3) img_pca = (img_pca - img_pca.min())/(img_pca.max() - img_pca.min()) return img_pca.astype(np.float32), label
之后数据集映射时换成parse_image_pca即可。
方案3:换用支持多通道的预训练模型
找针对遥感图像预训练的模型(比如在SENet、ResNet基础上用遥感数据集预训练的版本),或者直接从头训练模型。但这种方法会丢失ImageNet预训练权重的优势,训练时间和成本更高。
内容的提问来源于stack exchange,提问作者Brocoli
相关产品推荐
相关产品推荐

