如何基于预定义掩码分割图像并高效提取网格字母子图
解决方案
核心思路
用Numpy掩码快速定位有效区域(去除空白),再通过数组分块高效分割5x5网格,全程避免冗余循环,把时间复杂度降到O(w*h)(仅遍历图像一次),适合低功耗设备和大规模处理。
步骤1:批量去除图像边界空白
先把每张图像的周围空白裁掉,得到仅包含5x5网格的规整图像:
import cv2 import numpy as np import os from PIL import Image import pytesseract def crop_blank_borders(img): # 转灰度图,二值化(背景设为0,文字设为255) gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) _, binary = cv2.threshold(gray, 240, 255, cv2.THRESH_BINARY_INV) # 用Numpy掩码找非空白区域的行和列 non_zero_rows = np.where(binary.sum(axis=1) > 0)[0] non_zero_cols = np.where(binary.sum(axis=0) > 0)[0] # 裁剪到有效区域 cropped = img[non_zero_rows[0]:non_zero_rows[-1]+1, non_zero_cols[0]:non_zero_cols[-1]+1] return cropped
步骤2:高效分割5x5网格
裁剪后的图像是规整的,直接用Numpy分块,避免循环遍历,这一步是O(1)时间复杂度(因为是创建数组视图,不是复制数据):
def split_5x5_grid(img): # 把图像分成5行,每行再分成5列 rows = np.array_split(img, 5, axis=0) grid = [] for row in rows: cols = np.array_split(row, 5, axis=1) grid.extend(cols) # 返回25个子图的列表 return grid
步骤3:子图处理(去内部空白+统一尺寸)
对每个子图,同样用掩码裁剪内部空白,再resize到统一尺寸:
def process_subimg(subimg, target_size=(64, 64)): # 裁剪子图内部空白 gray = cv2.cvtColor(subimg, cv2.COLOR_BGR2GRAY) _, binary = cv2.threshold(gray, 240, 255, cv2.THRESH_BINARY_INV) non_zero_rows = np.where(binary.sum(axis=1) > 0)[0] non_zero_cols = np.where(binary.sum(axis=0) > 0)[0] if len(non_zero_rows) == 0 or len(non_zero_cols) == 0: return None # 跳过空白子图 cropped_sub = subimg[non_zero_rows[0]:non_zero_rows[-1]+1, non_zero_cols[0]:non_zero_cols[-1]+1] # 统一尺寸 resized = cv2.resize(cropped_sub, target_size, interpolation=cv2.INTER_AREA) return resized
步骤4:批量处理并保存
整合以上步骤,遍历所有图像,提取子图并保存到对应文件夹:
def batch_process(image_paths, target_size=(64, 64)): for idx, img_path in enumerate(image_paths, start=1): # 读取图像 img = cv2.imread(img_path) if img is None: continue # 1. 裁剪边界空白 cropped_img = crop_blank_borders(img) # 2. 分割5x5网格 subimgs = split_5x5_grid(cropped_img) # 3. 处理每个子图并保存 for subimg in subimgs: processed = process_subimg(subimg, target_size) if processed is None: continue # 识别子图中的字母(这里用pytesseract,可根据需求替换) text = pytesseract.image_to_string(processed, config='--psm 10 --oem 3 -c tessedit_char_whitelist=ABCDEFGHIJKLMNOPQRSTUVWXYZ').strip().upper() if not text: continue # 创建字母文件夹 save_dir = os.path.join(os.getcwd(), text) os.makedirs(save_dir, exist_ok=True) # 保存图像,命名为"序号.jpg" save_path = os.path.join(save_dir, f"{idx}.jpg") cv2.imwrite(save_path, processed) # 示例调用:替换为你的图像路径列表 image_paths = [f"image_{i}.jpg" for i in range(1, 101)] batch_process(image_paths)
关键优化点
- 用Numpy的向量化操作替代循环,避免O(n²)的冗余计算,大幅降低CPU占用
- 数组分块用
array_split创建视图,不复制数据,内存效率更高 - 掩码裁剪仅遍历图像一次,比硬编码定位更灵活,适配不同的空白边界
内容的提问来源于stack exchange,提问作者Matic
相关产品推荐
相关产品推荐

