如何利用Google Cloud Vision OCR获取手写段落的边界框?
解决Google Cloud Vision手写识别的段落边界框与块容器问题
一、在Document Text Detection中提取段落边界框
Document Text Detection API本身已包含段落的边界框信息,只需从响应的层级结构中提取即可。API返回的Page对象下,blocks包含多个文本块,每个block的paragraphs列表中,每个Paragraph元素都带有bounding_box字段,该字段存储了段落的边界框顶点坐标。
Python代码示例
from google.cloud import vision_v1 def get_paragraph_bounding_boxes(image_path): client = vision_v1.ImageAnnotatorClient() with open(image_path, "rb") as image_file: content = image_file.read() image = vision_v1.Image(content=content) response = client.document_text_detection(image=image) pages = response.full_text_annotation.pages paragraph_data = [] for page in pages: for block in page.blocks: for paragraph in block.paragraphs: # 拼接段落文本 paragraph_text = "".join([word.text for word in paragraph.words]) # 提取边界框顶点坐标 vertices = [{"x": vertex.x, "y": vertex.y} for vertex in paragraph.bounding_box.vertices] paragraph_data.append({ "text": paragraph_text, "bounding_box": vertices }) if response.error.message: raise Exception(f"API调用错误: {response.error.message}") return paragraph_data # 使用示例 result = get_paragraph_bounding_boxes("handwritten_note.jpg") for item in result: print(f"段落文本: {item['text']}") print(f"边界框坐标: {item['bounding_box']}\n")
二、在Text Detection中模拟块/段落容器
Text Detection API返回的text_annotations中,第一个元素是全文,后续元素为独立单词(带各自边界框)。可通过分析单词的位置坐标,将空间相邻的单词聚类为段落/块——核心逻辑是基于单词的垂直间距设定阈值,判断是否属于同一段落。
Python代码示例
from google.cloud import vision_v1 def cluster_words_into_paragraphs(image_path, vertical_threshold=20): client = vision_v1.ImageAnnotatorClient() with open(image_path, "rb") as image_file: content = image_file.read() image = vision_v1.Image(content=content) response = client.text_detection(image=image) annotations = response.text_annotations if len(annotations) < 2: return [] # 提取单个单词及其位置信息 words = [] for ann in annotations[1:]: y_min = min([v.y for v in ann.bounding_poly.vertices]) y_max = max([v.y for v in ann.bounding_poly.vertices]) y_center = (y_min + y_max) / 2 words.append({ "text": ann.description, "y_center": y_center, "bounding_box": [{"x": v.x, "y": v.y} for v in ann.bounding_poly.vertices] }) # 按垂直位置排序,保证从上到下处理 words.sort(key=lambda x: x["y_center"]) paragraphs = [] current_paragraph = { "text": words[0]["text"], "bounding_box": words[0]["bounding_box"], "y_center": words[0]["y_center"] } for word in words[1:]: # 判断当前单词与段落的垂直距离是否在阈值内 distance = abs(word["y_center"] - current_paragraph["y_center"]) if distance <= vertical_threshold: # 合并到当前段落,更新文本和边界框 current_paragraph["text"] += " " + word["text"] # 计算合并后的边界框(取所有顶点的最小/最大坐标) curr_x_min = min(v["x"] for v in current_paragraph["bounding_box"]) curr_x_max = max(v["x"] for v in current_paragraph["bounding_box"]) curr_y_min = min(v["y"] for v in current_paragraph["bounding_box"]) curr_y_max = max(v["y"] for v in current_paragraph["bounding_box"]) word_x_min = min(v["x"] for v in word["bounding_box"]) word_x_max = max(v["x"] for v in word["bounding_box"]) word_y_min = min(v["y"] for v in word["bounding_box"]) word_y_max = max(v["y"] for v in word["bounding_box"]) current_paragraph["bounding_box"] = [ {"x": min(curr_x_min, word_x_min), "y": min(curr_y_min, word_y_min)}, {"x": max(curr_x_max, word_x_max), "y": min(curr_y_min, word_y_min)}, {"x": max(curr_x_max, word_x_max), "y": max(curr_y_max, word_y_max)}, {"x": min(curr_x_min, word_x_min), "y": max(curr_y_max, word_y_max)} ] current_paragraph["y_center"] = word["y_center"] else: # 启动新段落 paragraphs.append(current_paragraph) current_paragraph = { "text": word["text"], "bounding_box": word["bounding_box"], "y_center": word["y_center"] } # 添加最后一个段落 paragraphs.append(current_paragraph) if response.error.message: raise Exception(f"API调用错误: {response.error.message}") return paragraphs # 使用示例,vertical_threshold可根据手写字体大小调整 result = cluster_words_into_paragraphs("handwritten_note.jpg", vertical_threshold=25) for idx, para in enumerate(result): print(f"段落{idx+1}文本: {para['text']}") print(f"段落{idx+1}边界框: {para['bounding_box']}\n")
内容的提问来源于stack exchange,提问作者Sam Rogers
相关产品推荐
相关产品推荐

