如何堆叠遥感图像数组并转换为指定维度以进行模型推理?
解决方案:遥感图像合并与样本生成适配模型输入
1. 统一所有图像尺寸
现有图像尺寸不一致,需先将所有图像resize到相同基础尺寸(建议选128的倍数,方便后续裁剪;例如(2560, 1920),128的倍数可保证无重叠裁剪时的样本数量计算更规整)。
from PIL import Image import numpy as np # 加载图像(假设已通过遥感工具导出为本地文件) img1 = np.array(Image.open("img1.tif")) # 维度(2583,1900) img2 = np.array(Image.open("img2.tif")) # 维度(2411,2571) img3 = np.array(Image.open("img3.tif")) # 维度(2583,1900) img4 = np.array(Image.open("img4.tif")) # 维度(2583,1900,3) # 定义目标基础尺寸(PIL的resize参数为(width, height),对应numpy的(height, width)) target_size = (1920, 2560) # 统一尺寸函数 def resize_img(arr, target_size): img = Image.fromarray(arr) resized_img = img.resize(target_size, Image.Resampling.LANCZOS) return np.array(resized_img) # 执行尺寸统一 img1_resized = resize_img(img1, target_size) img2_resized = resize_img(img2, target_size) img3_resized = resize_img(img3, target_size) img4_resized = resize_img(img4, target_size)
2. 扩展单通道图像的通道维度
三个二维单通道图像需扩展为三维格式(增加通道维度),才能与三通道图像在通道维度合并:
# 为单通道图像添加通道轴 img1_expanded = np.expand_dims(img1_resized, axis=-1) # 维度变为(2560, 1920, 1) img2_expanded = np.expand_dims(img2_resized, axis=-1) img3_expanded = np.expand_dims(img3_resized, axis=-1)
3. 通道维度合并所有图像
将四个图像在通道维度拼接,得到6通道的合并图像,匹配模型输入的通道数要求:
# 通道维度合并,最终维度为(2560, 1920, 6) merged_img = np.concatenate([img1_expanded, img2_expanded, img3_expanded, img4_resized], axis=-1)
4. 生成128×128的训练样本(目标3039个)
通过两种方式生成满足数量要求的样本:
方式1:滑动窗口+随机裁剪组合
先通过滑动窗口生成批量样本,再用随机裁剪补充不足的数量:
def extract_sliding_patches(arr, patch_size=128, stride=64): patches = [] h, w, _ = arr.shape # 滑动窗口遍历 for i in range(0, h - patch_size + 1, stride): for j in range(0, w - patch_size + 1, stride): patch = arr[i:i+patch_size, j:j+patch_size, :] patches.append(np.expand_dims(patch, axis=0)) # 扩展为(1,128,128,6)格式 return np.concatenate(patches, axis=0) # 步长64生成重叠样本,约可得到39×29=1131个样本 sliding_patches = extract_sliding_patches(merged_img) # 随机裁剪补充剩余样本 def random_crop_samples(arr, patch_size=128, need_num=3039): existing_num = len(arr) patches = [] h, w, _ = arr.shape while len(patches) + existing_num < need_num: # 随机生成裁剪左上角坐标 i = np.random.randint(0, h - patch_size + 1) j = np.random.randint(0, w - patch_size + 1) patch = arr[i:i+patch_size, j:j+patch_size, :] patches.append(np.expand_dims(patch, axis=0)) return np.concatenate([arr] + patches, axis=0)[:need_num] # 生成最终3039个样本 final_samples = random_crop_samples(sliding_patches) # final_samples.shape = (3039, 128, 128, 6),单个样本取索引即可得到(1,128,128,6)格式
方式2:TensorFlow批量提取补丁(高效适合训练场景)
若使用TensorFlow训练模型,可直接用内置函数批量生成:
import tensorflow as tf # 转换为Tensor并添加batch维度 merged_tensor = tf.expand_dims(tf.convert_to_tensor(merged_img), axis=0) # 批量提取补丁 patches_tensor = tf.image.extract_patches( images=merged_tensor, sizes=[1, 128, 128, 1], strides=[1, 64, 64, 1], rates=[1, 1, 1, 1], padding='VALID' ) # 重塑为样本格式 patches_reshaped = tf.reshape(patches_tensor, [-1, 128, 128, 6]) # 若数量不足,补充随机裁剪样本 remaining_num = 3039 - tf.shape(patches_reshaped)[0] if remaining_num > 0: random_patches = tf.image.random_crop(merged_tensor, [remaining_num, 128, 128, 6]) final_samples = tf.concat([patches_reshaped, random_patches], axis=0)
关键注意事项
- 确保所有图像数值范围一致,可统一归一化到
[0,1]或[-1,1],避免模型训练波动。 - 若原图像尺寸与目标尺寸差异过大,可先对原图像做分块裁剪,再统一通道后合并,减少resize带来的细节丢失。
- 随机裁剪时可加入旋转、翻转等数据增强,同时提升样本多样性和数量。
内容的提问来源于stack exchange,提问作者Kaushik Vezzu
相关产品推荐
相关产品推荐

