如何从DocumentAI OCR响应中提取块的行并添加标签与缩进标识
解决DocumentAI块内行访问问题并实现编号与缩进标注
问题根源
你遇到的block['lines']访问报错,通常是两种原因:
- 若使用DocumentAI客户端库返回的Protobuf对象,需用属性访问而非字典索引(即
block.lines而非block['lines']) - 若解析的是JSON响应,需确认字段名是否为
lines(注意大小写匹配DocumentAI JSON响应的字段格式)
实现方案
以下是完整代码示例,涵盖两种响应格式的处理,同时实现cluster编号、line编号及缩进特征识别:
1. 处理Protobuf响应(客户端库直接返回)
from google.cloud import documentai_v1 as documentai def process_documentai_response(document): page = document.pages[0] page_width = page.width # 获取页面宽度用于缩进计算 indent_threshold = page_width * 0.05 # 缩进阈值设为页面宽度的5% page_center = page_width / 2 for cluster_idx, block in enumerate(page.blocks, start=1): # 标记cluster编号 print(f"Cluster {cluster_idx}:") if not block.lines: print(" 无行数据") continue for line_idx, line in enumerate(block.lines, start=1): # 标记line编号 # 提取行边界框坐标 vertices = line.bounding_poly.vertices line_left = vertices[0].x line_right = vertices[2].x line_center = (line_left + line_right) / 2 # 判断缩进特征 indent_type = "无缩进" if line_left > indent_threshold: indent_type = "左缩进" elif abs(line_center - page_center) < indent_threshold: indent_type = "居中" elif (page_width - line_right) > indent_threshold: indent_type = "右缩进" # 提取行文本 line_text = "".join([token.text for token in line.tokens]) print(f" Line {line_idx}: {line_text} | 缩进类型: {indent_type}")
2. 处理JSON格式响应(已转成字典)
import json def process_documentai_json_response(json_doc): page = json_doc["pages"][0] page_width = page["width"] indent_threshold = page_width * 0.05 page_center = page_width / 2 for cluster_idx, block in enumerate(page["blocks"], start=1): print(f"Cluster {cluster_idx}:") # 检查lines字段是否存在 if "lines" not in block or not block["lines"]: print(" 无行数据") continue for line_idx, line in enumerate(block["lines"], start=1): vertices = line["boundingPoly"]["vertices"] line_left = vertices[0]["x"] line_right = vertices[2]["x"] line_center = (line_left + line_right) / 2 indent_type = "无缩进" if line_left > indent_threshold: indent_type = "左缩进" elif abs(line_center - page_center) < indent_threshold: indent_type = "居中" elif (page_width - line_right) > indent_threshold: indent_type = "右缩进" line_text = "".join([token["text"] for token in line["tokens"]]) print(f" Line {line_idx}: {line_text} | 缩进类型: {indent_type}")
关键注意事项
- 若仍报错,先打印
block的完整结构(print(block)或print(json.dumps(block, indent=2))),确认行数据所在的实际字段名 - 缩进阈值可根据手写信件的实际排版调整,避免误判
- 多页面场景需遍历
document.pages而非仅取第一页
内容的提问来源于stack exchange,提问作者adeshina Ibrahim
相关产品推荐
相关产品推荐

