You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用Google Cloud Vision OCR获取手写段落的边界框?

解决Google Cloud Vision手写识别的段落边界框与块容器问题

一、在Document Text Detection中提取段落边界框

Document Text Detection API本身已包含段落的边界框信息,只需从响应的层级结构中提取即可。API返回的Page对象下,blocks包含多个文本块,每个block的paragraphs列表中,每个Paragraph元素都带有bounding_box字段,该字段存储了段落的边界框顶点坐标。

Python代码示例

from google.cloud import vision_v1

def get_paragraph_bounding_boxes(image_path):
    client = vision_v1.ImageAnnotatorClient()

    with open(image_path, "rb") as image_file:
        content = image_file.read()

    image = vision_v1.Image(content=content)
    response = client.document_text_detection(image=image)
    pages = response.full_text_annotation.pages

    paragraph_data = []
    for page in pages:
        for block in page.blocks:
            for paragraph in block.paragraphs:
                # 拼接段落文本
                paragraph_text = "".join([word.text for word in paragraph.words])
                # 提取边界框顶点坐标
                vertices = [{"x": vertex.x, "y": vertex.y} for vertex in paragraph.bounding_box.vertices]
                paragraph_data.append({
                    "text": paragraph_text,
                    "bounding_box": vertices
                })

    if response.error.message:
        raise Exception(f"API调用错误: {response.error.message}")
    
    return paragraph_data

# 使用示例
result = get_paragraph_bounding_boxes("handwritten_note.jpg")
for item in result:
    print(f"段落文本: {item['text']}")
    print(f"边界框坐标: {item['bounding_box']}\n")

二、在Text Detection中模拟块/段落容器

Text Detection API返回的text_annotations中,第一个元素是全文,后续元素为独立单词(带各自边界框)。可通过分析单词的位置坐标,将空间相邻的单词聚类为段落/块——核心逻辑是基于单词的垂直间距设定阈值,判断是否属于同一段落。

Python代码示例

from google.cloud import vision_v1

def cluster_words_into_paragraphs(image_path, vertical_threshold=20):
    client = vision_v1.ImageAnnotatorClient()

    with open(image_path, "rb") as image_file:
        content = image_file.read()

    image = vision_v1.Image(content=content)
    response = client.text_detection(image=image)
    annotations = response.text_annotations

    if len(annotations) < 2:
        return []
    
    # 提取单个单词及其位置信息
    words = []
    for ann in annotations[1:]:
        y_min = min([v.y for v in ann.bounding_poly.vertices])
        y_max = max([v.y for v in ann.bounding_poly.vertices])
        y_center = (y_min + y_max) / 2
        words.append({
            "text": ann.description,
            "y_center": y_center,
            "bounding_box": [{"x": v.x, "y": v.y} for v in ann.bounding_poly.vertices]
        })
    
    # 按垂直位置排序,保证从上到下处理
    words.sort(key=lambda x: x["y_center"])

    paragraphs = []
    current_paragraph = {
        "text": words[0]["text"],
        "bounding_box": words[0]["bounding_box"],
        "y_center": words[0]["y_center"]
    }

    for word in words[1:]:
        # 判断当前单词与段落的垂直距离是否在阈值内
        distance = abs(word["y_center"] - current_paragraph["y_center"])
        if distance <= vertical_threshold:
            # 合并到当前段落,更新文本和边界框
            current_paragraph["text"] += " " + word["text"]
            # 计算合并后的边界框(取所有顶点的最小/最大坐标)
            curr_x_min = min(v["x"] for v in current_paragraph["bounding_box"])
            curr_x_max = max(v["x"] for v in current_paragraph["bounding_box"])
            curr_y_min = min(v["y"] for v in current_paragraph["bounding_box"])
            curr_y_max = max(v["y"] for v in current_paragraph["bounding_box"])
            
            word_x_min = min(v["x"] for v in word["bounding_box"])
            word_x_max = max(v["x"] for v in word["bounding_box"])
            word_y_min = min(v["y"] for v in word["bounding_box"])
            word_y_max = max(v["y"] for v in word["bounding_box"])
            
            current_paragraph["bounding_box"] = [
                {"x": min(curr_x_min, word_x_min), "y": min(curr_y_min, word_y_min)},
                {"x": max(curr_x_max, word_x_max), "y": min(curr_y_min, word_y_min)},
                {"x": max(curr_x_max, word_x_max), "y": max(curr_y_max, word_y_max)},
                {"x": min(curr_x_min, word_x_min), "y": max(curr_y_max, word_y_max)}
            ]
            current_paragraph["y_center"] = word["y_center"]
        else:
            # 启动新段落
            paragraphs.append(current_paragraph)
            current_paragraph = {
                "text": word["text"],
                "bounding_box": word["bounding_box"],
                "y_center": word["y_center"]
            }
    # 添加最后一个段落
    paragraphs.append(current_paragraph)

    if response.error.message:
        raise Exception(f"API调用错误: {response.error.message}")
    
    return paragraphs

# 使用示例,vertical_threshold可根据手写字体大小调整
result = cluster_words_into_paragraphs("handwritten_note.jpg", vertical_threshold=25)
for idx, para in enumerate(result):
    print(f"段落{idx+1}文本: {para['text']}")
    print(f"段落{idx+1}边界框: {para['bounding_box']}\n")

内容的提问来源于stack exchange,提问作者Sam Rogers

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 13:14:55