Python实现Azure OCR邻近文本边界框(Bounding Box)合并需求问询
合并Azure OCR邻近文本边界框的解决方案
嘿,我明白你现在的需求——把OCR返回的一个个小单词框合并成几个大的文本块框对吧?刚好我之前也做过类似的需求,给你分享一个实用的实现思路和代码修改方案:
核心思路
我们可以通过判断框与框之间的距离阈值来分组,把距离足够近的框归为同一组,然后对每组计算出一个覆盖所有子框的大边界框。具体步骤:
- 定义一个判断两个框是否“邻近”的规则(比如垂直/水平方向的间距小于设定阈值)
- 遍历所有框,把邻近的框合并为一个集群
- 对每个集群计算合并后的边界框(取所有子框的最小X、最小Y,最大X、最大Y)
修改后的完整代码
首先确保你导入了numpy(你的代码里最后用到了np.array,但开头没导入,记得加上),然后加入合并逻辑:
import requests %matplotlib inline import matplotlib.pyplot as plt from matplotlib.patches import Rectangle from PIL import Image from io import BytesIO import numpy as np # 新增导入 # Replace <Subscription Key> with your valid subscription key. subscription_key = "f244aa59ad4f4c05be907b4e78b7c6da" assert subscription_key vision_base_url = "https://westcentralus.api.cognitive.microsoft.com/vision/v2.0/" ocr_url = vision_base_url + "ocr" # Set image_url to the URL of an image that you want to analyze. image_url = "https://cdn-ayb.akinon.net/cms/2019/04/04/e494dce0-1e80-47eb-96c9-448960a71260.jpg" headers = {'Ocp-Apim-Subscription-Key': subscription_key} params = {'language': 'unk', 'detectOrientation': 'true'} data = {'url': image_url} response = requests.post(ocr_url, headers=headers, params=params, json=data) response.raise_for_status() analysis = response.json() # Extract the word bounding boxes and text. line_infos = [region["lines"] for region in analysis["regions"]] word_infos = [] for line in line_infos: for word_metadata in line: for word_info in word_metadata["words"]: word_infos.append(word_info) # ---------------------- 新增:合并邻近边界框的逻辑 ---------------------- def are_boxes_close(box1, box2, threshold=50): """判断两个框是否邻近,threshold是你设置的距离阈值(像素)""" # box格式:[Ymin, Xmin, Ymax, Xmax] y1_min, x1_min, y1_max, x1_max = box1 y2_min, x2_min, y2_max, x2_max = box2 # 计算两个框在垂直和水平方向的最小间距 # 如果框有重叠,间距为负数,肯定算邻近 vertical_gap = max(0, min(y1_max, y2_max) - max(y1_min, y2_min)) horizontal_gap = max(0, min(x1_max, x2_max) - max(x1_min, x2_min)) # 如果垂直方向重叠/间距小,且水平方向也邻近,就算同一组 # 你可以根据自己的图片调整判断逻辑,比如只看垂直或水平方向 return (y1_max + threshold >= y2_min and y2_max + threshold >= y1_min) and \ (x1_max + threshold >= x2_min and x2_max + threshold >= x1_min) def merge_close_boxes(boxes, threshold=50): """合并所有邻近的框,返回合并后的大框列表""" merged = [] visited = [False] * len(boxes) for i in range(len(boxes)): if visited[i]: continue # 初始化当前集群的框 current_cluster = [boxes[i]] visited[i] = True # 找所有和当前框邻近的框 for j in range(i+1, len(boxes)): if not visited[j] and any(are_boxes_close(box, boxes[j], threshold) for box in current_cluster): current_cluster.append(boxes[j]) visited[j] = True # 计算当前集群的合并框 y_min = min(box[0] for box in current_cluster) x_min = min(box[1] for box in current_cluster) y_max = max(box[2] for box in current_cluster) x_max = max(box[3] for box in current_cluster) merged.append([y_min, x_min, y_max, x_max]) return merged # 处理你的边界框数组 texts_boxes = [] texts = [] for word in word_infos: bbox = [int(num) for num in word["boundingBox"].split(",")] text = word["text"] new_box = [bbox[1], bbox[0], bbox[1]+bbox[3], bbox[0]+bbox[2]] texts_boxes.append(new_box) texts.append(text) texts_boxes_np = np.array(texts_boxes) # 合并框,这里阈值设为50,你可以根据实际图片调整 merged_boxes = merge_close_boxes(texts_boxes_np, threshold=50) # ---------------------- 合并逻辑结束 ---------------------- # Display the image and overlay it with the merged bounding boxes. plt.figure(figsize=(100, 20)) image = Image.open(BytesIO(requests.get(image_url).content)) ax = plt.imshow(image) # 绘制合并后的大框(用蓝色区分) for box in merged_boxes: y_min, x_min, y_max, x_max = box origin = (x_min, y_min) width = x_max - x_min height = y_max - y_min patch = Rectangle(origin, width, height, fill=False, linewidth=5, color='blue') ax.axes.add_patch(patch) # 可选:保留原来的小框(红色),对比效果 # for word in word_infos: # bbox = [int(num) for num in word["boundingBox"].split(",")] # text = word["text"] # origin = (bbox[0], bbox[1]) # patch = Rectangle(origin, bbox[2], bbox[3], fill=False, linewidth=3, color='r') # ax.axes.add_patch(patch) # plt.text(origin[0], origin[1], text, fontsize=2, weight="bold", va="top") plt.axis("off") plt.show() # 打印合并后的框 print("合并后的边界框:") print(np.array(merged_boxes))
关键说明
- 阈值调整:
threshold=50是我给的参考值,你可以根据图片中文本块的间距调整——如果文本块之间距离大,就调大阈值;如果小就调小。 - 邻近判断逻辑:
are_boxes_close函数里的判断规则是“垂直和水平方向都足够近”,你也可以根据需求修改,比如只看垂直方向(适合竖排文本)或者只看水平方向(适合横排文本)。 - 合并效果:针对你提供的边界框数组,用这个代码应该会合并成2个大框,正好符合你的需求。
内容的提问来源于stack exchange,提问作者Baron
相关产品推荐
相关产品推荐

