GCP Document AI批量处理超50份PDF文件的实现方法咨询
解决GCP Document AI批量处理50份文档上限的分批次方案
针对你存储在Cloud Storage同一文件夹下的800份扫描版PDF,以下是几种按50份/请求分批次提交Document AI批量处理的可行方案:
方案一:Shell脚本+gcloud命令行手动分批次
适合快速验证或小批量场景,步骤如下:
- 列出目标文件夹下所有PDF并保存到本地文件:
gsutil ls gs://你的存储桶名称/目标文件夹路径/*.pdf > pdf_list.txt - 将文件列表按每50行分割成多个批次文件:
执行后会生成split -l 50 pdf_list.txt batch_batch_aa、batch_ab等包含50个PDF路径的文件。 - 循环提交每个批次的处理请求:
创建一个shell脚本(比如submit_batches.sh),内容如下:
赋予脚本执行权限并运行:#!/bin/bash PROCESSOR_ID="你的处理器ID" LOCATION="处理器所在区域(如us)" INPUT_BUCKET="gs://你的存储桶名称/目标文件夹路径/" OUTPUT_BUCKET="gs://你的存储桶名称/输出文件夹路径/" for batch_file in batch_*; do # 将批次文件中的路径转为逗号分隔的字符串 INPUT_URIS=$(paste -sd "," "$batch_file") echo "提交批次:$batch_file" gcloud document-ai processors batch-process documents \ --processor="$PROCESSOR_ID" \ --location="$LOCATION" \ --input-uris="$INPUT_URIS" \ --output-destination="$OUTPUT_BUCKET" donechmod +x submit_batches.sh && ./submit_batches.sh
方案二:Python脚本自动化分批次
适合800份这种较大规模的场景,可自动完成文件枚举、分批次、提交请求全流程:
- 先安装依赖包:
pip install google-cloud-documentai google-cloud-storage - 创建Python脚本(比如
batch_process_docs.py):from google.cloud import documentai_v1 as documentai from google.cloud import storage def get_all_pdf_uris(bucket_name, folder_prefix): """获取指定存储桶文件夹下所有PDF的GCS路径""" storage_client = storage.Client() bucket = storage_client.get_bucket(bucket_name) blobs = bucket.list_blobs(prefix=folder_prefix) return [f"gs://{bucket_name}/{blob.name}" for blob in blobs if blob.name.endswith(".pdf")] def submit_batch_request(processor_client, processor_full_name, input_uris, output_bucket, output_subfolder): """提交单个批次的Document AI处理请求""" # 配置输出路径 gcs_output_config = documentai.GcsOutputConfig( gcs_uri=f"gs://{output_bucket}/{output_subfolder}/" ) # 构造每个文档的输入配置 input_configs = [ documentai.BatchProcessRequest.BatchInputConfig( gcs_source=uri, mime_type="application/pdf" ) for uri in input_uris ] # 构造批量处理请求 request = documentai.BatchProcessRequest( name=processor_full_name, input_configs=input_configs, gcs_output_config=gcs_output_config ) # 发起请求并等待完成(可根据需求改为异步监控) operation = processor_client.batch_process_documents(request) print(f"批次处理中,操作ID:{operation.operation.name}") operation.result() print(f"批次处理完成,输出路径:gs://{output_bucket}/{output_subfolder}/") if __name__ == "__main__": # 替换为你的GCP资源信息 PROJECT_ID = "你的项目ID" LOCATION = "处理器所在区域(如us-central1)" PROCESSOR_ID = "你的处理器ID" BUCKET_NAME = "你的存储桶名称" PDF_FOLDER = "存储PDF的文件夹路径/" OUTPUT_FOLDER = "输出结果的根文件夹路径/" BATCH_SIZE = 50 # 初始化客户端 processor_client = documentai.DocumentProcessorServiceClient() processor_full_name = processor_client.processor_path(PROJECT_ID, LOCATION, PROCESSOR_ID) # 获取所有PDF路径 all_pdf_uris = get_all_pdf_uris(BUCKET_NAME, PDF_FOLDER) total_batches = (len(all_pdf_uris) + BATCH_SIZE -1) // BATCH_SIZE print(f"共检测到{len(all_pdf_uris)}份PDF,将分为{total_batches}批次处理") # 分批次提交请求 for batch_idx in range(total_batches): start_idx = batch_idx * BATCH_SIZE end_idx = start_idx + BATCH_SIZE batch_uris = all_pdf_uris[start_idx:end_idx] output_subfolder = f"{OUTPUT_FOLDER}/batch_{batch_idx+1}" print(f"开始处理第{batch_idx+1}/{total_batches}批次,共{len(batch_uris)}份文档") submit_batch_request(processor_client, processor_full_name, batch_uris, BUCKET_NAME, output_subfolder) - 替换脚本中的配置参数后运行即可。
方案三:Cloud Functions自动化触发处理
如果后续还有新增PDF需要处理,可配置完全自动化的流程:
- 创建Cloud Function,设置触发条件为Cloud Storage对象创建(针对目标PDF文件夹)或定期触发(比如每天扫描一次文件夹)
- 在函数中实现逻辑:
- 扫描目标文件夹,筛选出未处理的PDF(可通过记录已处理文件路径到Firestore或存储桶的标记文件实现)
- 按50份为一组,自动提交Document AI批量处理请求
- 标记已处理的文件,避免重复处理
注意事项
- 确保执行操作的服务账号拥有
roles/documentai.processorUser权限,以及Cloud Storage的读写权限 - 单个批量处理请求的总文件大小不能超过5GB,若单PDF文件较大,需适当调整批次大小
- 可通过GCP控制台的Document AI页面查看每个批次的处理状态和结果
内容的提问来源于stack exchange,提问作者JMxnuell
相关产品推荐
相关产品推荐

