You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Document AI自定义提取器的JSON输出中提取有效信息?

解决Google Document AI自定义提取器冗余数据问题

问题背景

我正在使用Document AI的简单自定义提取器,尝试从上传的PDF中提取Country、Nombre、Address、Mail、City等字段。使用以下代码提取信息并打印输出JSON:

from google.cloud import documentai_v1 as documentai
import json
import os
from google.colab import files

# Credentials setup (assuming you've uploaded the service account key)
uploaded = files.upload()
key_file = list(uploaded.keys())[0]
os.environ["GOOGLE_APPLICATION_CREDENTIALS"] = key_file

# Configuration
PROJECT_ID = "682656916911"  # Replace with your project ID
LOCATION = "eu"  # Use the correct region
PROCESSOR_ID = "da26d6ce1aa73a53"  # Replace with your processor ID
DOCUMENT_PATH = "/content/W-8BEN.pdf"  # Path to your document

# Client setup
client_options = {"api_endpoint": "eu-documentai.googleapis.com"}
client = documentai.DocumentProcessorServiceClient(client_options=client_options)

# Request preparation
name = f"projects/{PROJECT_ID}/locations/{LOCATION}/processors/{PROCESSOR_ID}"
with open(DOCUMENT_PATH, "rb") as document_file:
    document_content = document_file.read()

request = {
    "name": name,
    "raw_document": {
        "content": document_content,
        "mime_type": "application/pdf"
    }
}

# Process document
response = client.process_document(request=request)

# Print response and extracted text
print(f"Response type: {type(response)}")
print(f"Document type: {type(response.document)}")

if document := response.document:  # Use walrus operator for cleaner assignment
    print("\nExtracted text:")
    print(document.text)

    # Convert Document object to dictionary (avoiding DESCRIPTOR field)
    document_dict = documentai.Document.to_dict(document)
    print("\nJSON representation of extracted data (excluding DESCRIPTOR):")
    print(json.dumps(document_dict, indent=4))

但输出的JSON包含大量坐标、布局等冗余信息,完全无法获取需要的字段值,示例片段如下:

{
    "layout": {
        "text_anchor": {
            "text_segments": [
                {
                    "start_index": "4392",
                    "end_index": "4396"
                }
            ],
            "content": ""
        },
        "confidence": 0.98868006,
        "bounding_poly": {
            "vertices": [
                {
                    "x": 522,
                    "y": 1972
                },
                {
                    "x": 549,
                    "y": 1972
                },
                {
                    "x": 549,
                    "y": 1994
                },
                {
                    "x": 522,
                    "y": 1994
                }
            ],
            "normalized_vertices": [
                {
                    "x": 0.29692832,
                    "y": 0.8668132
                },
                {
                    "x": 0.31228667,
                    "y": 0.8668132
                },
                {
                    "x": 0.31228667,
                    "y": 0.8764835
                },
                {
                    "x": 0.29692832,
                    "y": 0.8764835
                }
            ]
        },
        "orientation": 1
    },
    "detected_break": {
        "type_": 1
    },
    "detected_languages": [
        {
            "language_code": "en",
            "confidence": 1.0
        }
    ]
}

我需要过滤出所需字段的键值对,减少冗余信息,以便验证上传的PDF内容。


解决方案

自定义提取器的结果存储在response.document.entities中,直接遍历这个列表即可获取字段名和对应值,无需处理整个文档的冗余数据。

修改后的代码

将原代码中打印文档内容和完整JSON的部分替换为以下逻辑:

# Process document
response = client.process_document(request=request)

# Extract only custom fields from entities
if document := response.document:
    # Define the fields you need to extract
    target_fields = {"Country", "Nombre", "Address", "Mail", "City"}
    
    extracted_data = {}
    for entity in document.entities:
        # Get the custom field name and extracted value
        field_name = entity.type_
        field_value = entity.mention_text
        
        # Only keep fields in your target list
        if field_name in target_fields:
            extracted_data[field_name] = field_value
    
    # Print clean key-value pairs
    print("\nExtracted Key-Value Pairs:")
    print(json.dumps(extracted_data, indent=4))

代码说明

  • document.entities:自定义提取器识别出的所有字段都在这里,每个entity包含type_(你定义的字段名)和mention_text(提取到的实际值)
  • target_fields:设置你需要提取的字段集合,过滤掉无关字段
  • 最终输出的extracted_data是干净的键值对JSON,完全符合验证PDF内容的需求

内容的提问来源于stack exchange,提问作者Javier Romero Garcia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.15 22:47:14