Document AI在EC2实例中运行极慢甚至超时问题求助
Google Document AI EC2环境调用超时问题排查与优化方案
问题描述
本地环境调用Google Document AI的process_request接口处理2页约60KB的PDF文件,耗时稳定在8-20秒;但部署在美国EC2实例的预发布环境中,相同调用频繁超时,仅成功的一次耗时达34秒,无法满足业务性能要求。
相关代码实现
class DocumentAIClient: logger = get_logger(__name__) def __init__( self, project_id: str, location: str, processor_id: str, ): """ Initialize a Document AI client Args: project_id: Google Cloud project ID location: Location of the processor (e.g., "us") processor_id: Document AI processor ID """ os.environ["GOOGLE_APPLICATION_CREDENTIALS"] = os.environ[ "GOOGLE_APPLICATION_CREDENTIALS_FILE" ] self.project_id = project_id self.location = location self.processor_id = processor_id opts = ClientOptions(api_endpoint=f"{location}-documentai.googleapis.com") self.client = documentai.DocumentProcessorServiceClient(client_options=opts) def fetch_document_from_xxxxx( self, file_content: bytes, mime_type: str = "application/pdf", processor_version_id: str | None = None, ) -> documentai.Document: """ Process a document using Google Document AI Args: file_content: Binary content of the file mime_type: MIME type of the document field_mask: Optional fields to return in the Document object processor_version_id: Optional processor version ID pages: Optional list of specific pages to process Returns: Tuple containing the processed document and processing time """ start_time = time.time() if processor_version_id: name = self.client.processor_version_path( self.project_id, self.location, self.processor_id, processor_version_id ) else: name = self.client.processor_path( self.project_id, self.location, self.processor_id ) raw_document = documentai.RawDocument(content=file_content, mime_type=mime_type) process_options = documentai.ProcessOptions( individual_page_selector=documentai.ProcessOptions.IndividualPageSelector( pages=[1, 2] ) ) request = documentai.ProcessRequest( name=name, raw_document=raw_document, field_mask="text,entities,pages.pageNumber", process_options=process_options, ) try: result = self.client.process_document(request=request, timeout=90.0) except Exception as e: self.logger.exception("Error getting result: %s", e) raise end_time = time.time() processed_timing = end_time - start_time document = result.document self.logger.info( "Document AI processed document in %s seconds", processed_timing ) return document def fetch_pdf_from_s3( self, pdf_s3_url: str, ) -> bytes: self.logger = self.logger.bind(pdf_s3_url=pdf_s3_url) self.logger.info("S3 URL: %s", pdf_s3_url) parsed = urlparse(pdf_s3_url) key = parsed.path.lstrip("/") try: with default_storage.open(key, "rb") as file: data = file.read() self.logger.info("Fetched %s bytes from path %s", len(data), key) return data except Exception as e: self.logger.exception("Error fetching file from storage: %s", e) raise
优化方案
网络与客户端配置优化
- 验证EC2实例网络环境:确保EC2所在VPC带宽充足,无额外限流规则。若通过公网访问Google Cloud,建议切换为VPC对等连接或Cloud NAT,降低公网传输的延迟与不稳定性。
- 添加智能重试机制:针对网络波动导致的超时,集成Google官方推荐的重试策略(如
google.api_core.retry.Retry),对特定异常(如DeadlineExceeded)进行自动重试,避免单次超时导致请求失败。 - 启用HTTP/2传输:手动配置Document AI客户端启用HTTP/2协议,减少TCP连接建立的开销,提升数据传输效率。
API调用策略调整
- 对齐区域部署:确保Document AI处理器所在区域(如
us)与EC2实例区域一致,跨区域调用会显著增加网络延迟。 - 切换异步处理:改用
process_document_async接口提交处理任务,后续通过轮询获取结果,避免同步调用长时间阻塞,同时更好应对服务端负载波动。 - 精简请求配置:确认
process_options中的pages=[1,2]是否必要,若目标PDF固定为2页,可省略该配置,减少服务端额外处理逻辑。
代码细节优化
- 统一凭证配置:将
GOOGLE_APPLICATION_CREDENTIALS的设置移至应用启动阶段,避免每次实例化DocumentAIClient时重复修改环境变量。 - 细化性能日志:在EC2环境中添加DNS解析、连接建立、数据传输各阶段的耗时日志,精准定位性能瓶颈环节。
内容的提问来源于stack exchange,提问作者Wesley Bennett Loo Dela Cruz
相关产品推荐
相关产品推荐

