如何用Python检测含完整背景的文本框及其坐标?
解决方案
可以结合OpenCV的轮廓检测与easyocr的文本定位,精准找到包含完整背景的文本框,步骤如下:
核心思路
- 利用OpenCV的二值化处理,将深色背景与浅色文本分离,突出背景区域的轮廓;
- 通过easyocr获取所有可信文本的坐标范围;
- 在轮廓中筛选出能完全包含所有文本区域的背景轮廓,以此得到完整的背景矩形;
- 若轮廓检测效果不佳,可基于文本区域向外扩展,通过颜色判断找到背景边界。
代码实现
import cv2 import easyocr import numpy as np # 读取目标图片 img = cv2.imread('00025.jpg') gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) # 反二值化:将浅色文本转为黑色,深色背景转为白色,突出背景轮廓 _, binary = cv2.threshold(gray, 127, 255, cv2.THRESH_BINARY_INV) # 寻找图片中的所有外部轮廓 contours, _ = cv2.findContours(binary, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) # 用easyocr获取置信度高于0.75的文本框坐标 reader = easyocr.Reader(['en']) result = reader.readtext('00025.jpg') text_boxes = [res[0] for res in result if res[2] > 0.75] # 计算所有文本框的整体范围(最小外接矩形) all_points = [] for box in text_boxes: all_points.extend(box) all_points = np.array(all_points) x_min, y_min = np.min(all_points, axis=0) x_max, y_max = np.max(all_points, axis=0) # 筛选能完全包含文本范围的最大背景轮廓 target_contour = None max_area = 0 for cnt in contours: cnt_x, cnt_y, cnt_w, cnt_h = cv2.boundingRect(cnt) # 判断文本范围是否完全在当前轮廓内 if cnt_x <= x_min and cnt_y <= y_min and (cnt_x + cnt_w) >= x_max and (cnt_y + cnt_h) >= y_max: cnt_area = cnt_w * cnt_h if cnt_area > max_area: max_area = cnt_area target_contour = cnt # 获取最终的矩形坐标(格式与原代码一致:bottom_left, bottom_right, top_left, top_right) if target_contour is not None: x, y, w, h = cv2.boundingRect(target_contour) final_coords = [[x, y + h], [x + w, y + h], [x, y], [x + w, y]] print("最终背景文本框坐标:", final_coords) # 可选:绘制结果并保存 cv2.rectangle(img, (x, y), (x + w, y + h), (0, 255, 0), 2) cv2.imwrite('detected_result.jpg', img) else: # 若未找到合适轮廓,采用基于颜色的扩展方案 expand_step = 15 # 获取背景色样本(取文本区域外的深色区域) bg_color = gray[int(y_min)-10, int(x_min)-10] if y_min > 10 and x_min > 10 else 0 # 向外扩展直到遇到非背景色 while True: new_x_min = x_min - expand_step new_y_min = y_min - expand_step new_x_max = x_max + expand_step new_y_max = y_max + expand_step # 避免超出图片边界 if new_x_min < 0 or new_y_min < 0 or new_x_max >= img.shape[1] or new_y_max >= img.shape[0]: break # 检查扩展区域的颜色是否接近背景色 sample_color = gray[new_y_min, new_x_min] if abs(sample_color - bg_color) < 20: x_min, y_min, x_max, y_max = new_x_min, new_y_min, new_x_max, new_y_max else: break final_coords = [[x_min, y_max], [x_max, y_max], [x_min, y_min], [x_max, y_min]] print("最终背景文本框坐标(扩展版):", final_coords)
说明
- 轮廓检测优先:因为你的目标文本位于一个规整的深色背景块上,轮廓检测能直接定位到这个背景块的边界;
- 颜色扩展备选:如果图片背景复杂,轮廓检测失效,通过颜色判断向外扩展的方式可以适配更多场景;
- 可根据实际图片调整二值化阈值、扩展步长、颜色差值等参数,优化检测效果。
内容的提问来源于stack exchange,提问作者ale13
相关产品推荐
相关产品推荐

