You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现Azure OCR邻近文本边界框(Bounding Box)合并需求问询

合并Azure OCR邻近文本边界框的解决方案

嘿,我明白你现在的需求——把OCR返回的一个个小单词框合并成几个大的文本块框对吧?刚好我之前也做过类似的需求,给你分享一个实用的实现思路和代码修改方案:

核心思路

我们可以通过判断框与框之间的距离阈值来分组,把距离足够近的框归为同一组,然后对每组计算出一个覆盖所有子框的大边界框。具体步骤:

  1. 定义一个判断两个框是否“邻近”的规则(比如垂直/水平方向的间距小于设定阈值)
  2. 遍历所有框,把邻近的框合并为一个集群
  3. 对每个集群计算合并后的边界框(取所有子框的最小X、最小Y,最大X、最大Y)

修改后的完整代码

首先确保你导入了numpy(你的代码里最后用到了np.array,但开头没导入,记得加上),然后加入合并逻辑:

import requests
%matplotlib inline
import matplotlib.pyplot as plt
from matplotlib.patches import Rectangle
from PIL import Image
from io import BytesIO
import numpy as np  # 新增导入

# Replace <Subscription Key> with your valid subscription key.
subscription_key = "f244aa59ad4f4c05be907b4e78b7c6da"
assert subscription_key
vision_base_url = "https://westcentralus.api.cognitive.microsoft.com/vision/v2.0/"
ocr_url = vision_base_url + "ocr"

# Set image_url to the URL of an image that you want to analyze.
image_url = "https://cdn-ayb.akinon.net/cms/2019/04/04/e494dce0-1e80-47eb-96c9-448960a71260.jpg"

headers = {'Ocp-Apim-Subscription-Key': subscription_key}
params = {'language': 'unk', 'detectOrientation': 'true'}
data = {'url': image_url}
response = requests.post(ocr_url, headers=headers, params=params, json=data)
response.raise_for_status()
analysis = response.json()

# Extract the word bounding boxes and text.
line_infos = [region["lines"] for region in analysis["regions"]]
word_infos = []
for line in line_infos:
    for word_metadata in line:
        for word_info in word_metadata["words"]:
            word_infos.append(word_info)

# ---------------------- 新增:合并邻近边界框的逻辑 ----------------------
def are_boxes_close(box1, box2, threshold=50):
    """判断两个框是否邻近,threshold是你设置的距离阈值(像素)"""
    # box格式:[Ymin, Xmin, Ymax, Xmax]
    y1_min, x1_min, y1_max, x1_max = box1
    y2_min, x2_min, y2_max, x2_max = box2

    # 计算两个框在垂直和水平方向的最小间距
    # 如果框有重叠,间距为负数,肯定算邻近
    vertical_gap = max(0, min(y1_max, y2_max) - max(y1_min, y2_min))
    horizontal_gap = max(0, min(x1_max, x2_max) - max(x1_min, x2_min))
    
    # 如果垂直方向重叠/间距小,且水平方向也邻近,就算同一组
    # 你可以根据自己的图片调整判断逻辑,比如只看垂直或水平方向
    return (y1_max + threshold >= y2_min and y2_max + threshold >= y1_min) and \
           (x1_max + threshold >= x2_min and x2_max + threshold >= x1_min)

def merge_close_boxes(boxes, threshold=50):
    """合并所有邻近的框,返回合并后的大框列表"""
    merged = []
    visited = [False] * len(boxes)

    for i in range(len(boxes)):
        if visited[i]:
            continue
        # 初始化当前集群的框
        current_cluster = [boxes[i]]
        visited[i] = True
        # 找所有和当前框邻近的框
        for j in range(i+1, len(boxes)):
            if not visited[j] and any(are_boxes_close(box, boxes[j], threshold) for box in current_cluster):
                current_cluster.append(boxes[j])
                visited[j] = True
        # 计算当前集群的合并框
        y_min = min(box[0] for box in current_cluster)
        x_min = min(box[1] for box in current_cluster)
        y_max = max(box[2] for box in current_cluster)
        x_max = max(box[3] for box in current_cluster)
        merged.append([y_min, x_min, y_max, x_max])
    
    return merged

# 处理你的边界框数组
texts_boxes = []
texts = []
for word in word_infos:
    bbox = [int(num) for num in word["boundingBox"].split(",")]
    text = word["text"]
    new_box = [bbox[1], bbox[0], bbox[1]+bbox[3], bbox[0]+bbox[2]]
    texts_boxes.append(new_box)
    texts.append(text)

texts_boxes_np = np.array(texts_boxes)
# 合并框,这里阈值设为50,你可以根据实际图片调整
merged_boxes = merge_close_boxes(texts_boxes_np, threshold=50)
# ---------------------- 合并逻辑结束 ----------------------

# Display the image and overlay it with the merged bounding boxes.
plt.figure(figsize=(100, 20))
image = Image.open(BytesIO(requests.get(image_url).content))
ax = plt.imshow(image)

# 绘制合并后的大框(用蓝色区分)
for box in merged_boxes:
    y_min, x_min, y_max, x_max = box
    origin = (x_min, y_min)
    width = x_max - x_min
    height = y_max - y_min
    patch = Rectangle(origin, width, height, fill=False, linewidth=5, color='blue')
    ax.axes.add_patch(patch)

# 可选:保留原来的小框(红色),对比效果
# for word in word_infos:
#     bbox = [int(num) for num in word["boundingBox"].split(",")]
#     text = word["text"]
#     origin = (bbox[0], bbox[1])
#     patch = Rectangle(origin, bbox[2], bbox[3], fill=False, linewidth=3, color='r')
#     ax.axes.add_patch(patch)
#     plt.text(origin[0], origin[1], text, fontsize=2, weight="bold", va="top")

plt.axis("off")
plt.show()

# 打印合并后的框
print("合并后的边界框:")
print(np.array(merged_boxes))

关键说明

  1. 阈值调整:threshold=50是我给的参考值,你可以根据图片中文本块的间距调整——如果文本块之间距离大,就调大阈值;如果小就调小。
  2. 邻近判断逻辑:are_boxes_close函数里的判断规则是“垂直和水平方向都足够近”,你也可以根据需求修改,比如只看垂直方向(适合竖排文本)或者只看水平方向(适合横排文本)。
  3. 合并效果:针对你提供的边界框数组,用这个代码应该会合并成2个大框,正好符合你的需求。

内容的提问来源于stack exchange,提问作者Baron

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:17:39