You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GCP Document AI批量处理超50份PDF文件的实现方法咨询

解决GCP Document AI批量处理50份文档上限的分批次方案

针对你存储在Cloud Storage同一文件夹下的800份扫描版PDF,以下是几种按50份/请求分批次提交Document AI批量处理的可行方案:

方案一:Shell脚本+gcloud命令行手动分批次

适合快速验证或小批量场景,步骤如下:

  1. 列出目标文件夹下所有PDF并保存到本地文件:
    gsutil ls gs://你的存储桶名称/目标文件夹路径/*.pdf > pdf_list.txt
    
  2. 将文件列表按每50行分割成多个批次文件:
    split -l 50 pdf_list.txt batch_
    
    执行后会生成batch_aa、batch_ab等包含50个PDF路径的文件。
  3. 循环提交每个批次的处理请求:
    创建一个shell脚本(比如submit_batches.sh),内容如下:
    #!/bin/bash
    PROCESSOR_ID="你的处理器ID"
    LOCATION="处理器所在区域(如us)"
    INPUT_BUCKET="gs://你的存储桶名称/目标文件夹路径/"
    OUTPUT_BUCKET="gs://你的存储桶名称/输出文件夹路径/"
    
    for batch_file in batch_*; do
        # 将批次文件中的路径转为逗号分隔的字符串
        INPUT_URIS=$(paste -sd "," "$batch_file")
        echo "提交批次:$batch_file"
        gcloud document-ai processors batch-process documents \
            --processor="$PROCESSOR_ID" \
            --location="$LOCATION" \
            --input-uris="$INPUT_URIS" \
            --output-destination="$OUTPUT_BUCKET"
    done
    
    赋予脚本执行权限并运行:chmod +x submit_batches.sh && ./submit_batches.sh

方案二:Python脚本自动化分批次

适合800份这种较大规模的场景,可自动完成文件枚举、分批次、提交请求全流程:

  1. 先安装依赖包:
    pip install google-cloud-documentai google-cloud-storage
    
  2. 创建Python脚本(比如batch_process_docs.py):
    from google.cloud import documentai_v1 as documentai
    from google.cloud import storage
    
    def get_all_pdf_uris(bucket_name, folder_prefix):
        """获取指定存储桶文件夹下所有PDF的GCS路径"""
        storage_client = storage.Client()
        bucket = storage_client.get_bucket(bucket_name)
        blobs = bucket.list_blobs(prefix=folder_prefix)
        return [f"gs://{bucket_name}/{blob.name}" for blob in blobs if blob.name.endswith(".pdf")]
    
    def submit_batch_request(processor_client, processor_full_name, input_uris, output_bucket, output_subfolder):
        """提交单个批次的Document AI处理请求"""
        # 配置输出路径
        gcs_output_config = documentai.GcsOutputConfig(
            gcs_uri=f"gs://{output_bucket}/{output_subfolder}/"
        )
        # 构造每个文档的输入配置
        input_configs = [
            documentai.BatchProcessRequest.BatchInputConfig(
                gcs_source=uri, mime_type="application/pdf"
            ) for uri in input_uris
        ]
        # 构造批量处理请求
        request = documentai.BatchProcessRequest(
            name=processor_full_name,
            input_configs=input_configs,
            gcs_output_config=gcs_output_config
        )
        # 发起请求并等待完成(可根据需求改为异步监控)
        operation = processor_client.batch_process_documents(request)
        print(f"批次处理中,操作ID:{operation.operation.name}")
        operation.result()
        print(f"批次处理完成,输出路径:gs://{output_bucket}/{output_subfolder}/")
    
    if __name__ == "__main__":
        # 替换为你的GCP资源信息
        PROJECT_ID = "你的项目ID"
        LOCATION = "处理器所在区域(如us-central1)"
        PROCESSOR_ID = "你的处理器ID"
        BUCKET_NAME = "你的存储桶名称"
        PDF_FOLDER = "存储PDF的文件夹路径/"
        OUTPUT_FOLDER = "输出结果的根文件夹路径/"
        BATCH_SIZE = 50
    
        # 初始化客户端
        processor_client = documentai.DocumentProcessorServiceClient()
        processor_full_name = processor_client.processor_path(PROJECT_ID, LOCATION, PROCESSOR_ID)
    
        # 获取所有PDF路径
        all_pdf_uris = get_all_pdf_uris(BUCKET_NAME, PDF_FOLDER)
        total_batches = (len(all_pdf_uris) + BATCH_SIZE -1) // BATCH_SIZE
        print(f"共检测到{len(all_pdf_uris)}份PDF,将分为{total_batches}批次处理")
    
        # 分批次提交请求
        for batch_idx in range(total_batches):
            start_idx = batch_idx * BATCH_SIZE
            end_idx = start_idx + BATCH_SIZE
            batch_uris = all_pdf_uris[start_idx:end_idx]
            output_subfolder = f"{OUTPUT_FOLDER}/batch_{batch_idx+1}"
            print(f"开始处理第{batch_idx+1}/{total_batches}批次,共{len(batch_uris)}份文档")
            submit_batch_request(processor_client, processor_full_name, batch_uris, BUCKET_NAME, output_subfolder)
    
  3. 替换脚本中的配置参数后运行即可。

方案三:Cloud Functions自动化触发处理

如果后续还有新增PDF需要处理,可配置完全自动化的流程:

  • 创建Cloud Function,设置触发条件为Cloud Storage对象创建(针对目标PDF文件夹)或定期触发(比如每天扫描一次文件夹)
  • 在函数中实现逻辑:
    1. 扫描目标文件夹,筛选出未处理的PDF(可通过记录已处理文件路径到Firestore或存储桶的标记文件实现)
    2. 按50份为一组,自动提交Document AI批量处理请求
    3. 标记已处理的文件,避免重复处理

注意事项

  • 确保执行操作的服务账号拥有roles/documentai.processorUser权限,以及Cloud Storage的读写权限
  • 单个批量处理请求的总文件大小不能超过5GB,若单PDF文件较大,需适当调整批次大小
  • 可通过GCP控制台的Document AI页面查看每个批次的处理状态和结果

内容的提问来源于stack exchange,提问作者JMxnuell

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 02:48:11