You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Document AI Python客户端本地批量处理PDF文件?

问题分析

需要批量处理本地多文件夹下的PDF文档(包含原生及扫描件,部分文档页数超过15页),使用Google Document AI 2.20.1版本时,尝试构建批量处理逻辑遇到AttributeError: module 'google.cloud.documentai' has no attribute 'types'报错,且官方示例多依赖Google Cloud Storage(GCS),希望实现本地提交批量处理或接受GCS方案。

解决方案

一、修复types属性报错

Document AI v1版本(如2.20.1)已将类型定义直接整合到documentai模块下,无需通过types子模块访问:

  • 错误写法:documentai.types.Document
  • 正确写法:documentai.Document

所有原documentai.types.*的类型,直接使用documentai.*即可调用。

二、本地批量提交文件到Document AI(非完全离线)

注意:Document AI是云端服务,不存在完全离线的本地处理能力,此处的“本地批量处理”指从本地读取文件后,批量提交到云端API处理。对于页数超过15页的文档,必须使用批量处理API(而非同步单文件API)。

实现步骤

  1. 遍历本地文件夹,收集所有PDF文件路径:
import os

def get_all_pdf_paths(root_dir):
    pdf_paths = []
    for dirpath, _, filenames in os.walk(root_dir):
        for filename in filenames:
            if filename.lower().endswith('.pdf'):
                pdf_paths.append(os.path.join(dirpath, filename))
    return pdf_paths

root_directory = "/path/to/your/pdf/folders"
file_paths = get_all_pdf_paths(root_directory)
  1. 构建批量处理请求并提交:
from google.api_core.client_options import ClientOptions
from google.cloud import documentai
from google.cloud.documentai_v1.types import batch_process
from google.cloud import storage
import json

# 配置客户端
project_id = "your-project-id"
location = "us"  # 替换为你的区域,如eu、asia-southeast1等
processor_id = "your-processor-id"
processor_version = "rc"  # 或指定具体版本号
bucket_name = "your-gcs-bucket-name"  # 批量处理结果需存储到GCS

client_options = ClientOptions(api_endpoint=f"{location}-documentai.googleapis.com")
client = documentai.DocumentProcessorServiceClient(client_options=client_options)

# 构建批量请求
request = batch_process.BatchProcessRequest(
    name=client.processor_path(project_id, location, processor_id),
    input_configs=[],
    output_config=batch_process.OutputConfig(
        gcs_output_config=batch_process.GcsOutputConfig(
            uri=f"gs://{bucket_name}/documentai-output/"
        )
    )
)

# 逐个添加本地文件到请求
for file_path in file_paths:
    with open(file_path, "rb") as f:
        file_content = f.read()
    
    input_config = batch_process.BatchProcessRequest.BatchInputConfig(
        raw_document=documentai.RawDocument(
            content=file_content,
            mime_type="application/pdf"
        ),
        document_id=os.path.basename(file_path)  # 用文件名关联后续结果
    )
    request.input_configs.append(input_config)

# 发送批量处理请求
operation = client.batch_process_documents(request)
print("等待批量处理完成...")
response = operation.result(timeout=3600)  # 根据文件数量调整超时时间

# 下载并处理结果
storage_client = storage.Client()
bucket = storage_client.bucket(bucket_name)

for result in response.output_configs:
    gcs_uri = result.gcs_output_config.uri
    blob_path = gcs_uri.replace(f"gs://{bucket_name}/", "")
    blob = bucket.blob(blob_path)
    result_content = blob.download_as_bytes()
    
    # 解析结果文档
    document = documentai.Document.from_json(result_content)
    # 关联原文件路径
    original_filename = document.document_id
    original_path = next(p for p in file_paths if os.path.basename(p) == original_filename)
    
    # 保存提取的文本
    with open(f"{original_filename}_extracted.txt", "w", encoding="utf-8") as f:
        f.write(document.text)
    # 保存元数据
    metadata = {
        "original_path": original_path,
        "total_pages": len(document.pages),
        "mime_type": document.mime_type
    }
    with open(f"{original_filename}_metadata.json", "w") as f:
        json.dump(metadata, f, indent=2)

三、推荐方案:使用GCS进行批量处理(更稳定高效)

对于大量文件(尤其是页数超过15页的文档),官方推荐使用GCS存储源文件和处理结果,避免本地提交大文件的超时风险:

实现步骤

  1. 上传本地PDF文件到GCS:
def upload_to_gcs(file_paths, bucket_name):
    storage_client = storage.Client()
    bucket = storage_client.bucket(bucket_name)
    for file_path in file_paths:
        blob_name = os.path.basename(file_path)
        blob = bucket.blob(blob_name)
        blob.upload_from_filename(file_path)
        print(f"已上传 {file_path} 到 gs://{bucket_name}/{blob_name}")

upload_to_gcs(file_paths, bucket_name)
  1. 构建GCS源的批量处理请求:
request = batch_process.BatchProcessRequest(
    name=client.processor_path(project_id, location, processor_id),
    input_configs=[
        batch_process.BatchProcessRequest.BatchInputConfig(
            gcs_source=batch_process.GcsSource(
                uri=f"gs://{bucket_name}/*.pdf"
            ),
            mime_type="application/pdf"
        )
    ],
    output_config=batch_process.OutputConfig(
        gcs_output_config=batch_process.GcsOutputConfig(
            uri=f"gs://{bucket_name}/documentai-output/"
        )
    )
)

operation = client.batch_process_documents(request)
response = operation.result(timeout=3600)
  1. 下载并关联结果(代码同本地提交方案的结果处理部分)

关键注意事项

  • Document AI所有处理逻辑均在云端执行,本地仅负责提交请求和接收结果,无完全离线处理能力。
  • 批量处理API强制要求结果输出到GCS,无论源文件来自本地还是GCS。
  • 确保已配置Google Cloud本地认证(如执行gcloud auth application-default login),且账号拥有Document AI和Storage的对应权限。
  • 单文件超过1GB时,优先使用GCS上传,避免本地字节流提交的超时问题。

内容的提问来源于stack exchange,提问作者Vojta Partík

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 15:49:54