You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Document AI将JSON文件直接保存到GCS根目录而非子文件夹?

解决Document AI解析结果直接保存到GCS存储桶根目录的方法

Document AI批量处理默认会在目标存储桶生成多层子目录(如YYYYMMDDHHMMSS/<文档ID>/结构),要直接保存到根目录,可根据你的文档规模选择以下两种方案:

方案1:使用同步处理(适合单小文件)

同步处理(ProcessDocument API)支持直接指定输出文件的完整GCS路径,跳过默认层级结构。需注意同步处理有单文档大小限制(最大20MB,页数上限依处理器类型而定,如OCR处理器最多100页)。

示例代码(Python)

from google.cloud import documentai_v1 as documentai

def process_document_to_root_gcs(project_id, location, processor_id, input_gcs_uri, output_gcs_uri):
    client = documentai.DocumentProcessorServiceClient()
    name = client.processor_path(project_id, location, processor_id)

    # 配置输入源
    input_config = documentai.DocumentProcessorService.ProcessRequest.InputConfig(
        gcs_source=documentai.GcsSource(uri=input_gcs_uri),
        mime_type="application/pdf"
    )

    # 配置输出到GCS根目录的指定文件
    output_config = documentai.DocumentProcessorService.ProcessRequest.OutputConfig(
        gcs_destination=documentai.GcsDestination(uri=output_gcs_uri)
    )

    request = documentai.DocumentProcessorService.ProcessRequest(
        name=name,
        input_config=input_config,
        output_config=output_config
    )

    response = client.process_document(request=request)
    print(f"解析结果已保存到: {output_gcs_uri}")

# 调用示例
process_document_to_root_gcs(
    project_id="your-project-id",
    location="us",
    processor_id="your-processor-id",
    input_gcs_uri="gs://source-bucket/input.pdf",
    output_gcs_uri="gs://target-bucket/output.json"
)

方案2:批量处理后自动移动文件(适合大量文件)

如果必须用批量处理(支持大文件/多文档批量处理),可通过Cloud Function监听目标存储桶的对象创建事件,自动将子目录中的JSON文件移动到根目录,并清理原层级文件夹。

实现逻辑

  1. 创建Cloud Function,触发类型选择Cloud Storage > 对象创建,指定目标存储桶。
  2. 在函数中:
    • 捕获新生成的JSON文件路径(如gs://target-bucket/20240520123456/doc-123/output.json)
    • 提取文件名(如output.json,或自定义为原PDF文件名+.json)
    • 将文件移动到根目录路径(gs://target-bucket/output.json)
    • 删除原层级的文件夹及文件

示例函数代码(Python)

import os
from google.cloud import storage

def move_json_to_root(event, context):
    storage_client = storage.Client()
    bucket_name = event['bucket']
    file_path = event['name']
    
    # 仅处理Document AI生成的JSON文件(路径包含层级结构)
    if not file_path.endswith('.json') or '/' not in file_path:
        return
    
    # 提取目标文件名(可自定义,比如用原PDF文件名替换)
    target_filename = os.path.basename(file_path)
    
    # 移动文件到根目录
    bucket = storage_client.bucket(bucket_name)
    source_blob = bucket.blob(file_path)
    destination_blob = bucket.blob(target_filename)
    
    # 复制并删除原文件
    destination_blob.copy_from(source_blob)
    source_blob.delete()
    
    # 删除空的父文件夹(可选)
    parent_folder = os.path.dirname(file_path)
    blobs = list(bucket.list_blobs(prefix=parent_folder+'/'))
    if len(blobs) == 0:
        bucket.blob(parent_folder+'/').delete()
    
    print(f"已将 {file_path} 移动到根目录: {target_filename}")

注意事项

  • 确保Cloud Function拥有GCS存储桶的读写权限(配置合适的IAM角色,如roles/storage.objectAdmin)。
  • 若存在重名文件,需添加文件名冲突处理逻辑(如添加时间戳后缀)。

内容的提问来源于stack exchange,提问作者c0nfusion

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 00:03:09