You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于预定义掩码分割图像并高效提取网格字母子图

解决方案

核心思路

用Numpy掩码快速定位有效区域(去除空白),再通过数组分块高效分割5x5网格,全程避免冗余循环,把时间复杂度降到O(w*h)(仅遍历图像一次),适合低功耗设备和大规模处理。

步骤1:批量去除图像边界空白

先把每张图像的周围空白裁掉,得到仅包含5x5网格的规整图像:

import cv2
import numpy as np
import os
from PIL import Image
import pytesseract

def crop_blank_borders(img):
    # 转灰度图,二值化(背景设为0,文字设为255)
    gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
    _, binary = cv2.threshold(gray, 240, 255, cv2.THRESH_BINARY_INV)
    
    # 用Numpy掩码找非空白区域的行和列
    non_zero_rows = np.where(binary.sum(axis=1) > 0)[0]
    non_zero_cols = np.where(binary.sum(axis=0) > 0)[0]
    
    # 裁剪到有效区域
    cropped = img[non_zero_rows[0]:non_zero_rows[-1]+1, non_zero_cols[0]:non_zero_cols[-1]+1]
    return cropped

步骤2:高效分割5x5网格

裁剪后的图像是规整的,直接用Numpy分块,避免循环遍历,这一步是O(1)时间复杂度(因为是创建数组视图,不是复制数据):

def split_5x5_grid(img):
    # 把图像分成5行,每行再分成5列
    rows = np.array_split(img, 5, axis=0)
    grid = []
    for row in rows:
        cols = np.array_split(row, 5, axis=1)
        grid.extend(cols)
    # 返回25个子图的列表
    return grid

步骤3:子图处理(去内部空白+统一尺寸)

对每个子图,同样用掩码裁剪内部空白,再resize到统一尺寸:

def process_subimg(subimg, target_size=(64, 64)):
    # 裁剪子图内部空白
    gray = cv2.cvtColor(subimg, cv2.COLOR_BGR2GRAY)
    _, binary = cv2.threshold(gray, 240, 255, cv2.THRESH_BINARY_INV)
    non_zero_rows = np.where(binary.sum(axis=1) > 0)[0]
    non_zero_cols = np.where(binary.sum(axis=0) > 0)[0]
    if len(non_zero_rows) == 0 or len(non_zero_cols) == 0:
        return None  # 跳过空白子图
    cropped_sub = subimg[non_zero_rows[0]:non_zero_rows[-1]+1, non_zero_cols[0]:non_zero_cols[-1]+1]
    
    # 统一尺寸
    resized = cv2.resize(cropped_sub, target_size, interpolation=cv2.INTER_AREA)
    return resized

步骤4:批量处理并保存

整合以上步骤,遍历所有图像,提取子图并保存到对应文件夹:

def batch_process(image_paths, target_size=(64, 64)):
    for idx, img_path in enumerate(image_paths, start=1):
        # 读取图像
        img = cv2.imread(img_path)
        if img is None:
            continue
        
        # 1. 裁剪边界空白
        cropped_img = crop_blank_borders(img)
        # 2. 分割5x5网格
        subimgs = split_5x5_grid(cropped_img)
        
        # 3. 处理每个子图并保存
        for subimg in subimgs:
            processed = process_subimg(subimg, target_size)
            if processed is None:
                continue
            
            # 识别子图中的字母(这里用pytesseract,可根据需求替换)
            text = pytesseract.image_to_string(processed, config='--psm 10 --oem 3 -c tessedit_char_whitelist=ABCDEFGHIJKLMNOPQRSTUVWXYZ').strip().upper()
            if not text:
                continue
            
            # 创建字母文件夹
            save_dir = os.path.join(os.getcwd(), text)
            os.makedirs(save_dir, exist_ok=True)
            
            # 保存图像,命名为"序号.jpg"
            save_path = os.path.join(save_dir, f"{idx}.jpg")
            cv2.imwrite(save_path, processed)

# 示例调用:替换为你的图像路径列表
image_paths = [f"image_{i}.jpg" for i in range(1, 101)]
batch_process(image_paths)

关键优化点

  • 用Numpy的向量化操作替代循环,避免O(n²)的冗余计算,大幅降低CPU占用
  • 数组分块用array_split创建视图,不复制数据,内存效率更高
  • 掩码裁剪仅遍历图像一次,比硬编码定位更灵活,适配不同的空白边界

内容的提问来源于stack exchange,提问作者Matic

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 00:35:20