如何从Document AI自定义提取器的JSON输出中提取有效信息?
解决Google Document AI自定义提取器冗余数据问题
问题背景
我正在使用Document AI的简单自定义提取器,尝试从上传的PDF中提取Country、Nombre、Address、Mail、City等字段。使用以下代码提取信息并打印输出JSON:
from google.cloud import documentai_v1 as documentai import json import os from google.colab import files # Credentials setup (assuming you've uploaded the service account key) uploaded = files.upload() key_file = list(uploaded.keys())[0] os.environ["GOOGLE_APPLICATION_CREDENTIALS"] = key_file # Configuration PROJECT_ID = "682656916911" # Replace with your project ID LOCATION = "eu" # Use the correct region PROCESSOR_ID = "da26d6ce1aa73a53" # Replace with your processor ID DOCUMENT_PATH = "/content/W-8BEN.pdf" # Path to your document # Client setup client_options = {"api_endpoint": "eu-documentai.googleapis.com"} client = documentai.DocumentProcessorServiceClient(client_options=client_options) # Request preparation name = f"projects/{PROJECT_ID}/locations/{LOCATION}/processors/{PROCESSOR_ID}" with open(DOCUMENT_PATH, "rb") as document_file: document_content = document_file.read() request = { "name": name, "raw_document": { "content": document_content, "mime_type": "application/pdf" } } # Process document response = client.process_document(request=request) # Print response and extracted text print(f"Response type: {type(response)}") print(f"Document type: {type(response.document)}") if document := response.document: # Use walrus operator for cleaner assignment print("\nExtracted text:") print(document.text) # Convert Document object to dictionary (avoiding DESCRIPTOR field) document_dict = documentai.Document.to_dict(document) print("\nJSON representation of extracted data (excluding DESCRIPTOR):") print(json.dumps(document_dict, indent=4))
但输出的JSON包含大量坐标、布局等冗余信息,完全无法获取需要的字段值,示例片段如下:
{ "layout": { "text_anchor": { "text_segments": [ { "start_index": "4392", "end_index": "4396" } ], "content": "" }, "confidence": 0.98868006, "bounding_poly": { "vertices": [ { "x": 522, "y": 1972 }, { "x": 549, "y": 1972 }, { "x": 549, "y": 1994 }, { "x": 522, "y": 1994 } ], "normalized_vertices": [ { "x": 0.29692832, "y": 0.8668132 }, { "x": 0.31228667, "y": 0.8668132 }, { "x": 0.31228667, "y": 0.8764835 }, { "x": 0.29692832, "y": 0.8764835 } ] }, "orientation": 1 }, "detected_break": { "type_": 1 }, "detected_languages": [ { "language_code": "en", "confidence": 1.0 } ] }
我需要过滤出所需字段的键值对,减少冗余信息,以便验证上传的PDF内容。
解决方案
自定义提取器的结果存储在response.document.entities中,直接遍历这个列表即可获取字段名和对应值,无需处理整个文档的冗余数据。
修改后的代码
将原代码中打印文档内容和完整JSON的部分替换为以下逻辑:
# Process document response = client.process_document(request=request) # Extract only custom fields from entities if document := response.document: # Define the fields you need to extract target_fields = {"Country", "Nombre", "Address", "Mail", "City"} extracted_data = {} for entity in document.entities: # Get the custom field name and extracted value field_name = entity.type_ field_value = entity.mention_text # Only keep fields in your target list if field_name in target_fields: extracted_data[field_name] = field_value # Print clean key-value pairs print("\nExtracted Key-Value Pairs:") print(json.dumps(extracted_data, indent=4))
代码说明
document.entities:自定义提取器识别出的所有字段都在这里,每个entity包含type_(你定义的字段名)和mention_text(提取到的实际值)target_fields:设置你需要提取的字段集合,过滤掉无关字段- 最终输出的
extracted_data是干净的键值对JSON,完全符合验证PDF内容的需求
内容的提问来源于stack exchange,提问作者Javier Romero Garcia
相关产品推荐
相关产品推荐

