使用scikit-learn PatchExtractor后如何跟踪补丁所属原始图像?
跟踪scikit-learn PatchExtractor提取的补丁所属原始图像
这确实是个非常实际的痛点——当批量提取图像补丁后,没法溯源每个补丁的来源,后续的分析或者标注工作根本没法开展。我之前处理医学图像补丁的时候也遇到过一模一样的问题,下面分享几个亲测有效的解决方法:
方法一:基于补丁数量计算索引(适用于固定提取数量的场景)
如果你的PatchExtractor没有设置max_patches(即提取所有可能的补丁),或者每张图提取的补丁数量是固定的,那可以先计算单张图能生成的补丁数,再通过重复索引的方式生成对应关系。
举个具体的代码例子:
import numpy as np from sklearn.feature_extraction.image import PatchExtractor # 假设你的图像数组形状是(n_images, image_h, image_w) images = np.random.rand(10, 256, 256) # 示例:10张256x256的图像 patch_size = (64, 64) # 初始化提取器(默认步长等于patch_size,无重叠) extractor = PatchExtractor(patch_size=patch_size) # 计算单张图像能提取的补丁数量 n_patches_per_img = (images.shape[1] // patch_size[0]) * (images.shape[2] // patch_size[1]) # 生成每个补丁对应的原始图像索引:每张图对应n_patches_per_img个相同索引 patch_to_img_indices = np.repeat(np.arange(images.shape[0]), n_patches_per_img) # 提取补丁 patches = extractor.transform(images) # 现在patches[i]对应的原始图像索引就是patch_to_img_indices[i]
如果是有重叠的提取(需要设置step参数),那单张图的补丁数量计算要调整为:
step = (32, 32) # 步长设为32,重叠一半 n_patches_per_img = ((images.shape[1] - patch_size[0]) // step[0] + 1) * ((images.shape[2] - patch_size[1]) // step[1] + 1)
方法二:循环逐个提取并标记(通用场景,尤其适合随机提取)
如果你的提取器设置了max_patches(比如每张图随机提取N个补丁),这时候每张图的补丁数量不固定,方法一就失效了。这种情况下,最稳妥的方式是逐个处理每张图像,提取补丁的同时直接标记索引。
代码示例:
import numpy as np from sklearn.feature_extraction.image import PatchExtractor images = np.random.rand(10, 256, 256) patch_size = (64, 64) extractor = PatchExtractor(patch_size=patch_size, max_patches=10, random_state=42) # 每张图随机提10个补丁 patches_list = [] patch_indices_list = [] for img_idx, single_img in enumerate(images): # 注意:transform需要输入4D数组(批量),所以给单张图加一个维度 img_patches = extractor.transform(single_img[np.newaxis, ...]) patches_list.append(img_patches) # 给当前图像的所有补丁打上相同的索引 patch_indices_list.extend([img_idx] * len(img_patches)) # 合并所有补丁和索引 patches = np.concatenate(patches_list, axis=0) patch_to_img_indices = np.array(patch_indices_list)
这个方法不管你是固定数量还是随机提取,都能准确跟踪每个补丁的来源,唯一的小缺点是需要循环处理,但对于大多数图像数量来说,这点性能影响几乎可以忽略。
额外提示
如果你用的是skimage.util.extract_patches_2d(PatchExtractor内部其实也是调用这个函数),同样可以用上面的循环标记方法,逻辑完全一致。
内容的提问来源于stack exchange,提问作者John Smith
相关产品推荐
相关产品推荐

