如何高效生成包含原始图像坐标的图像重叠补丁?
提取图像补丁及对应坐标的实现方案
可以自己实现一个和sklearn.feature_extraction.image.extract_patches_2d输出顺序一致的函数,同时返回每个补丁在原始图像中的坐标信息,具体实现如下:
函数代码
import numpy as np def get_patches_and_coords_function(img, patch_dim=(40, 30)): img_height, img_width = img.shape[:2] patch_h, patch_w = patch_dim # 计算所有合法的起始坐标(y_top对应图像高度方向的行起始,x_top对应宽度方向的列起始) y_starts = np.arange(0, img_height - patch_h + 1, 1) x_starts = np.arange(0, img_width - patch_w + 1, 1) coords = [] patches_list = [] # 保持和sklearn工具输出的补丁顺序一致:先遍历行起始,再遍历列起始 for y_top in y_starts: for x_top in x_starts: # 提取当前补丁 patch = img[y_top:y_top+patch_h, x_top:x_top+patch_w] patches_list.append(patch) # 存储坐标信息 coords.append({ 'x_top': int(x_top), 'y_top': int(y_top), 'w': patch_w, 'h': patch_h }) # 将补丁列表转为numpy数组,和sklearn输出格式匹配 patches = np.stack(patches_list) return patches, coords
使用示例
from matplotlib import image # 加载图像为像素数组 img = image.imread('kolala.jpeg') # img.shape = (2000, 1000) # 提取补丁和对应坐标 patches, coords = get_patches_and_coords_function(img, patch_dim=(40, 30)) # 打印前3个坐标示例 for coord in coords[:3]: print(coord)
额外说明
- 函数输出的
patches数组格式和sklearn.extract_patches_2d完全一致,可直接替换使用 - 若需要自定义滑动步长,只需修改
np.arange中的步长参数即可,例如步长改为2时,写np.arange(0, img_height - patch_h +1, 2) - 坐标字典中的
x_top是补丁左上角的列索引,y_top是左上角的行索引,w和h为补丁的固定宽高
内容的提问来源于stack exchange,提问作者Becky W
相关产品推荐
相关产品推荐

