You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用OpenCV Python为文字绘制边界框并将裁剪图像保存到文件夹

实现方案

完整可运行代码

import cv2
import os
import matplotlib.pyplot as plt

# ---------------------- 配置参数 ----------------------
# 裁剪文字区域的保存目录
SAVE_CROP_DIR = "/content/cropped_words"
# 原图路径
IMG_PATH = "/content/passbook_64_0.jpeg"
# 带框结果图的保存路径
RESULT_IMG_PATH = "/content/result_with_boxes.jpg"
# ------------------------------------------------------

# 创建保存裁剪结果的文件夹,已存在则不报错
os.makedirs(SAVE_CROP_DIR, exist_ok=True)

# 读取原图,复制一份用于绘制边界框,避免破坏原始图像数据
origin_image = cv2.imread(IMG_PATH)
draw_image = origin_image.copy()
img_height, img_width = origin_image.shape[:2]

# 遍历所有检测结果,支持多页文档
for page_idx, page_coords in enumerate(words_abs_coords):
    for word_idx, (xmin, ymin, xmax, ymax) in enumerate(page_coords):
        # 边界校验,避免坐标超出图像范围导致报错
        xmin = max(0, xmin)
        ymin = max(0, ymin)
        xmax = min(img_width, xmax)
        ymax = min(img_height, ymax)

        # 绘制绿色边界框,线宽为2
        cv2.rectangle(draw_image, (xmin, ymin), (xmax, ymax), (0, 255, 0), 2)

        # 裁剪对应文字区域,OpenCV图像数组格式为 [y轴范围, x轴范围, 通道数]
        cropped_word = origin_image[ymin:ymax, xmin:xmax]
        # 按序号命名保存裁剪结果,避免重名覆盖
        crop_save_path = os.path.join(SAVE_CROP_DIR, f"page_{page_idx}_word_{word_idx}.jpg")
        cv2.imwrite(crop_save_path, cropped_word)

# 保存绘制完所有框的结果图
cv2.imwrite(RESULT_IMG_PATH, draw_image)

# 如需预览结果,需将OpenCV默认的BGR格式转为RGB格式适配matplotlib
plt.figure(figsize=(12, 12))
plt.imshow(cv2.cvtColor(draw_image, cv2.COLOR_BGR2RGB))
plt.axis("off")
plt.show()

关键逻辑说明

  • 批量绘制边界框:直接循环遍历所有提取到的文字坐标,重复调用cv2.rectangle即可在同一张图上叠加绘制所有框,无需每次生成新的图像对象。
  • 裁剪文字区域:OpenCV存储的图像是[y, x, channel]的三维数组,直接按坐标切片即可得到对应文字区域,注意y轴范围在前、x轴范围在后。
  • 批量存储:提前创建存储目录,按「页码+文字序号」的规则命名文件,可避免多页场景下的重名覆盖问题。

内容的提问来源于stack exchange,提问作者Asp Lab

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 17:54:10