如何让Google DocumentAI按行识别图片?OCR输出顺序异常求助
解决Google DocumentAI OCR文本顺序错乱问题
使用Google DocumentAI对图片进行OCR识别时,输出文本顺序完全错误。待识别图片本身已分段,本该按行读取,但识别结果先拆分图片再分别读取,导致顺序混乱。
待识别图片:
错误输出结果:
piece of clothing, = do up zip up something with difficulty. zip something up you fasten it using a zip. She zipped up the dress Hezipped his jeans up.
使用的Python代码:
def quickstart( project_id: str, location: str, processor_id: str, file_path: str, mime_type: str, processor_version_id: str = None ): # You must set the api_endpoint if you use a location other than 'us'. opts = ClientOptions(api_endpoint=f"{location}-documentai.googleapis.com") client = documentai.DocumentProcessorServiceClient(client_options=opts) # The full resource name of the processor, e.g.: # projects/project_id/locations/location/processor/processor_id # name = client.processor_path(project_id, location, processor_id) if processor_version_id: # The full resource name of the processor version, e.g.: # projects/{project_id}/locations/{location}/processors/{processor_id}/processorVersions/{processor_version_id} name = client.processor_version_path( project_id, location, processor_id, processor_version_id ) else: # The full resource name of the processor, e.g.: # projects/{project_id}/locations/{location}/processors/{processor_id} name = client.processor_path(project_id, location, processor_id) # Read the file into memory with open(file_path, "rb") as image: image_content = image.read() # Load Binary Data into Document AI RawDocument Object raw_document = documentai.RawDocument(content=image_content, mime_type=mime_type) # Configure the process request request = documentai.ProcessRequest(name=name, raw_document=raw_document) result = client.process_document(request=request) # For a full list of Document object attributes, please reference this page: # https://cloud.google.com/python/docs/reference/documentai/latest/google.cloud.documentai_v1.types.Document document = result.document # Read the text recognition output from the processor f.write(file_path + "\n") f.write(document.text)
问题原因
直接读取document.text会返回DocumentAI内部的识别顺序文本,而非页面的空间布局顺序,因此出现顺序错乱。
解决方案
通过遍历文档的页面、区块、段落和单词,利用其边界框坐标排序,还原页面的行读取顺序。修改后的代码如下:
def quickstart( project_id: str, location: str, processor_id: str, file_path: str, mime_type: str, processor_version_id: str = None ): opts = ClientOptions(api_endpoint=f"{location}-documentai.googleapis.com") client = documentai.DocumentProcessorServiceClient(client_options=opts) if processor_version_id: name = client.processor_version_path( project_id, location, processor_id, processor_version_id ) else: name = client.processor_path(project_id, location, processor_id) with open(file_path, "rb") as image: image_content = image.read() raw_document = documentai.RawDocument(content=image_content, mime_type=mime_type) request = documentai.ProcessRequest(name=name, raw_document=raw_document) result = client.process_document(request=request) document = result.document # 按空间顺序提取文本 sorted_text = [] for page in document.pages: # 按区块顶部y坐标排序,实现从上到下 blocks = sorted(page.blocks, key=lambda b: b.bounding_poly.normalized_vertices[0].y) for block in blocks: # 按段落顶部y坐标排序 paragraphs = sorted(block.paragraphs, key=lambda p: p.bounding_poly.normalized_vertices[0].y) for paragraph in paragraphs: # 按单词左侧x坐标排序,实现从左到右 words = sorted(paragraph.words, key=lambda w: w.bounding_poly.normalized_vertices[0].x) paragraph_text = "".join([word.text for word in words]) sorted_text.append(paragraph_text) # 拼接成最终文本 formatted_text = "\n".join(sorted_text) # 写入文件(替换为你的文件操作逻辑) with open("output.txt", "w", encoding="utf-8") as f: f.write(file_path + "\n") f.write(formatted_text)
代码说明
- 遍历文档每一页,对页面内的区块按顶部y坐标排序,确保从上到下读取;
- 区块内的段落同样按顶部y坐标排序;
- 段落内的单词按左侧x坐标排序,确保从左到右读取;
- 拼接排序后的文本,实现按行读取的效果。
内容的提问来源于stack exchange,提问作者jonah_w
相关产品推荐
相关产品推荐

