如何使用Document AI Python客户端本地批量处理PDF文件?
问题分析
需要批量处理本地多文件夹下的PDF文档(包含原生及扫描件,部分文档页数超过15页),使用Google Document AI 2.20.1版本时,尝试构建批量处理逻辑遇到AttributeError: module 'google.cloud.documentai' has no attribute 'types'报错,且官方示例多依赖Google Cloud Storage(GCS),希望实现本地提交批量处理或接受GCS方案。
解决方案
一、修复types属性报错
Document AI v1版本(如2.20.1)已将类型定义直接整合到documentai模块下,无需通过types子模块访问:
- 错误写法:
documentai.types.Document - 正确写法:
documentai.Document
所有原documentai.types.*的类型,直接使用documentai.*即可调用。
二、本地批量提交文件到Document AI(非完全离线)
注意:Document AI是云端服务,不存在完全离线的本地处理能力,此处的“本地批量处理”指从本地读取文件后,批量提交到云端API处理。对于页数超过15页的文档,必须使用批量处理API(而非同步单文件API)。
实现步骤
- 遍历本地文件夹,收集所有PDF文件路径:
import os def get_all_pdf_paths(root_dir): pdf_paths = [] for dirpath, _, filenames in os.walk(root_dir): for filename in filenames: if filename.lower().endswith('.pdf'): pdf_paths.append(os.path.join(dirpath, filename)) return pdf_paths root_directory = "/path/to/your/pdf/folders" file_paths = get_all_pdf_paths(root_directory)
- 构建批量处理请求并提交:
from google.api_core.client_options import ClientOptions from google.cloud import documentai from google.cloud.documentai_v1.types import batch_process from google.cloud import storage import json # 配置客户端 project_id = "your-project-id" location = "us" # 替换为你的区域,如eu、asia-southeast1等 processor_id = "your-processor-id" processor_version = "rc" # 或指定具体版本号 bucket_name = "your-gcs-bucket-name" # 批量处理结果需存储到GCS client_options = ClientOptions(api_endpoint=f"{location}-documentai.googleapis.com") client = documentai.DocumentProcessorServiceClient(client_options=client_options) # 构建批量请求 request = batch_process.BatchProcessRequest( name=client.processor_path(project_id, location, processor_id), input_configs=[], output_config=batch_process.OutputConfig( gcs_output_config=batch_process.GcsOutputConfig( uri=f"gs://{bucket_name}/documentai-output/" ) ) ) # 逐个添加本地文件到请求 for file_path in file_paths: with open(file_path, "rb") as f: file_content = f.read() input_config = batch_process.BatchProcessRequest.BatchInputConfig( raw_document=documentai.RawDocument( content=file_content, mime_type="application/pdf" ), document_id=os.path.basename(file_path) # 用文件名关联后续结果 ) request.input_configs.append(input_config) # 发送批量处理请求 operation = client.batch_process_documents(request) print("等待批量处理完成...") response = operation.result(timeout=3600) # 根据文件数量调整超时时间 # 下载并处理结果 storage_client = storage.Client() bucket = storage_client.bucket(bucket_name) for result in response.output_configs: gcs_uri = result.gcs_output_config.uri blob_path = gcs_uri.replace(f"gs://{bucket_name}/", "") blob = bucket.blob(blob_path) result_content = blob.download_as_bytes() # 解析结果文档 document = documentai.Document.from_json(result_content) # 关联原文件路径 original_filename = document.document_id original_path = next(p for p in file_paths if os.path.basename(p) == original_filename) # 保存提取的文本 with open(f"{original_filename}_extracted.txt", "w", encoding="utf-8") as f: f.write(document.text) # 保存元数据 metadata = { "original_path": original_path, "total_pages": len(document.pages), "mime_type": document.mime_type } with open(f"{original_filename}_metadata.json", "w") as f: json.dump(metadata, f, indent=2)
三、推荐方案:使用GCS进行批量处理(更稳定高效)
对于大量文件(尤其是页数超过15页的文档),官方推荐使用GCS存储源文件和处理结果,避免本地提交大文件的超时风险:
实现步骤
- 上传本地PDF文件到GCS:
def upload_to_gcs(file_paths, bucket_name): storage_client = storage.Client() bucket = storage_client.bucket(bucket_name) for file_path in file_paths: blob_name = os.path.basename(file_path) blob = bucket.blob(blob_name) blob.upload_from_filename(file_path) print(f"已上传 {file_path} 到 gs://{bucket_name}/{blob_name}") upload_to_gcs(file_paths, bucket_name)
- 构建GCS源的批量处理请求:
request = batch_process.BatchProcessRequest( name=client.processor_path(project_id, location, processor_id), input_configs=[ batch_process.BatchProcessRequest.BatchInputConfig( gcs_source=batch_process.GcsSource( uri=f"gs://{bucket_name}/*.pdf" ), mime_type="application/pdf" ) ], output_config=batch_process.OutputConfig( gcs_output_config=batch_process.GcsOutputConfig( uri=f"gs://{bucket_name}/documentai-output/" ) ) ) operation = client.batch_process_documents(request) response = operation.result(timeout=3600)
- 下载并关联结果(代码同本地提交方案的结果处理部分)
关键注意事项
- Document AI所有处理逻辑均在云端执行,本地仅负责提交请求和接收结果,无完全离线处理能力。
- 批量处理API强制要求结果输出到GCS,无论源文件来自本地还是GCS。
- 确保已配置Google Cloud本地认证(如执行
gcloud auth application-default login),且账号拥有Document AI和Storage的对应权限。 - 单文件超过1GB时,优先使用GCS上传,避免本地字节流提交的超时问题。
内容的提问来源于stack exchange,提问作者Vojta Partík
相关产品推荐
相关产品推荐

