You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Document AI在EC2实例中运行极慢甚至超时问题求助

Google Document AI EC2环境调用超时问题排查与优化方案

问题描述

本地环境调用Google Document AI的process_request接口处理2页约60KB的PDF文件,耗时稳定在8-20秒;但部署在美国EC2实例的预发布环境中,相同调用频繁超时,仅成功的一次耗时达34秒,无法满足业务性能要求。

相关代码实现

class DocumentAIClient:

    logger = get_logger(__name__)

    def __init__(
        self,
        project_id: str,
        location: str,
        processor_id: str,
    ):
        """
        Initialize a Document AI client

        Args:
            project_id: Google Cloud project ID
            location: Location of the processor (e.g., "us")
            processor_id: Document AI processor ID
        """

        os.environ["GOOGLE_APPLICATION_CREDENTIALS"] = os.environ[
            "GOOGLE_APPLICATION_CREDENTIALS_FILE"
        ]

        self.project_id = project_id
        self.location = location
        self.processor_id = processor_id
        opts = ClientOptions(api_endpoint=f"{location}-documentai.googleapis.com")
        self.client = documentai.DocumentProcessorServiceClient(client_options=opts)

    def fetch_document_from_xxxxx(
        self,
        file_content: bytes,
        mime_type: str = "application/pdf",
        processor_version_id: str | None = None,
    ) -> documentai.Document:
        """
        Process a document using Google Document AI

        Args:
            file_content: Binary content of the file
            mime_type: MIME type of the document
            field_mask: Optional fields to return in the Document object
            processor_version_id: Optional processor version ID
            pages: Optional list of specific pages to process

        Returns:
            Tuple containing the processed document and processing time
        """
        start_time = time.time()

        if processor_version_id:
            name = self.client.processor_version_path(
                self.project_id, self.location, self.processor_id, processor_version_id
            )
        else:
            name = self.client.processor_path(
                self.project_id, self.location, self.processor_id
            )
            
        raw_document = documentai.RawDocument(content=file_content, mime_type=mime_type)

        process_options = documentai.ProcessOptions(
            individual_page_selector=documentai.ProcessOptions.IndividualPageSelector(
                pages=[1, 2]
            )
        )

        request = documentai.ProcessRequest(
            name=name,
            raw_document=raw_document,
            field_mask="text,entities,pages.pageNumber",
            process_options=process_options,
        )

        try:
            result = self.client.process_document(request=request, timeout=90.0)
        except Exception as e:
            self.logger.exception("Error getting result: %s", e)
            raise
        
        end_time = time.time()
        processed_timing = end_time - start_time
        document = result.document
        self.logger.info(
            "Document AI processed document in %s seconds", processed_timing
        )

        return document

    def fetch_pdf_from_s3(
        self,
        pdf_s3_url: str,
    ) -> bytes:
        self.logger = self.logger.bind(pdf_s3_url=pdf_s3_url)
        self.logger.info("S3 URL: %s", pdf_s3_url)

        parsed = urlparse(pdf_s3_url)
        key = parsed.path.lstrip("/")

        try:
            with default_storage.open(key, "rb") as file:
                data = file.read()
            self.logger.info("Fetched %s bytes from path %s", len(data), key)
            return data
        except Exception as e:
            self.logger.exception("Error fetching file from storage: %s", e)
            raise

优化方案

网络与客户端配置优化

  • 验证EC2实例网络环境:确保EC2所在VPC带宽充足,无额外限流规则。若通过公网访问Google Cloud,建议切换为VPC对等连接或Cloud NAT,降低公网传输的延迟与不稳定性。
  • 添加智能重试机制:针对网络波动导致的超时,集成Google官方推荐的重试策略(如google.api_core.retry.Retry),对特定异常(如DeadlineExceeded)进行自动重试,避免单次超时导致请求失败。
  • 启用HTTP/2传输:手动配置Document AI客户端启用HTTP/2协议,减少TCP连接建立的开销,提升数据传输效率。

API调用策略调整

  • 对齐区域部署:确保Document AI处理器所在区域(如us)与EC2实例区域一致,跨区域调用会显著增加网络延迟。
  • 切换异步处理:改用process_document_async接口提交处理任务,后续通过轮询获取结果,避免同步调用长时间阻塞,同时更好应对服务端负载波动。
  • 精简请求配置:确认process_options中的pages=[1,2]是否必要,若目标PDF固定为2页,可省略该配置,减少服务端额外处理逻辑。

代码细节优化

  • 统一凭证配置:将GOOGLE_APPLICATION_CREDENTIALS的设置移至应用启动阶段,避免每次实例化DocumentAIClient时重复修改环境变量。
  • 细化性能日志:在EC2环境中添加DNS解析、连接建立、数据传输各阶段的耗时日志,精准定位性能瓶颈环节。

内容的提问来源于stack exchange,提问作者Wesley Bennett Loo Dela Cruz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 05:39:51