在Google Cloud Run实例中调用Google APIs出现请求超时问题排查
我通过以下脚本部署了Google Cloud Run实例:
gcloud run deploy <name-of-cloud-run> \ --image us-<region>-docker.pkg.dev/<artifact-location> \ --region <region> \ --service-account <service-account> \ --concurrency 50 \ --memory 1024Mi \ --min-instances 0 \ --max-instances 1 \ --port 5000 \ --timeout 1800 \ --env-vars-file .env.yaml \ --add-cloudsql-instances <cloud-sql-connection-string> \ --project <project>
该实例负责执行长时间运行的任务,遇到特定错误时会记录日志,任务完成时发布PubSub消息。但使用Google Cloud Python Libraries调用Cloud Logging和PubSub API时,均遇到相同错误:
Deadline of 60.0s exceeded while calling target function, last exception: 504 Deadline Exceeded
底层镜像的Dockerfile如下:
# Using python slim-buster FROM python:3.9-slim-buster # Setting the working directory to /app WORKDIR /app # Setting Shell SHELL [ "/bin/bash" , "-c" ] # Updating System RUN apt-get -y update # Installing Requirements COPY requirements.txt requirements.txt RUN pip3 install --upgrade pip && pip3 install -r requirements.txt # Copy the current directory contents into the container at /app COPY . /app ENTRYPOINT [ "gunicorn", "main:app" , "--worker-class", "gevent", "--workers", "8", "--max-requests", "10000", "--timeout", "1800" ] CMD [ "--bind", "0.0.0.0:5000" ]
此前类似部署未出现该问题,这次是首次遇到。参考相关回答认为是外部API问题,但我调用的是Google自身API,请问问题出在哪里?
1. Gevent协程与Google客户端库的兼容性问题
你使用了gevent作为gunicorn的worker类,但Google Cloud Python客户端库默认使用标准阻塞式IO,在gevent环境下如果未正确打补丁,会导致IO操作被阻塞,进而触发超时。
- 解决方法:在代码最开头添加gevent补丁,确保所有IO操作都被协程化:
注意:必须在导入任何其他Google客户端库或网络相关模块之前执行这行代码。from gevent import monkey monkey.patch_all()
2. Cloud Run实例资源限制
你设置了--max-instances 1和--concurrency 50,意味着单个实例要同时处理50个请求,而内存仅配置为1024Mi。当大量请求同时调用Google API时,可能因资源不足导致网络请求排队超时。
- 解决方法:
- 临时调高内存配置(比如改为
2048Mi)测试是否解决问题; - 降低并发数(比如改为
20),减少单个实例的负载; - 若任务允许,调高
max-instances让负载分散到多个实例。
- 临时调高内存配置(比如改为
3. Google客户端库的超时配置
默认情况下,Google Python客户端库的超时时间为60秒,若任务运行时间长或网络延迟高,这个阈值可能不够。
- 解决方法:初始化客户端时手动设置更长的超时时间,以PubSub为例:
Cloud Logging客户端同理,在初始化时指定超时参数。from google.cloud import pubsub_v1 publisher = pubsub_v1.PublisherClient( timeout=300 # 设置为300秒,根据实际需求调整 )
4. Cloud SQL连接的资源抢占
你添加了Cloud SQL实例连接,长时间运行的数据库操作可能占用大量资源,导致调用Logging/PubSub的网络请求无法获得足够带宽或CPU时间,进而超时。
- 解决方法:
- 检查数据库查询是否有优化空间(比如慢查询、未加索引);
- 考虑将数据库操作和API调用的逻辑异步分离,避免资源抢占。
5. Service Account权限与API访问限制
虽然调用的是Google自身API,但如果Service Account权限配置有误,或者API被限流,也可能表现为超时(实际是权限/限流错误被客户端包装成超时)。
- 解决方法:
- 确认Service Account拥有
roles/logging.logWriter和roles/pubsub.publisher权限; - 查看Cloud Logging和PubSub的API监控面板,检查是否有请求被限流的记录。
- 确认Service Account拥有
内容的提问来源于stack exchange,提问作者yashprkash

