Google Cloud Run上FastAPI+Uvicorn服务随机重启致交易中断问题
Cloud Run上FastAPI交易机器人随机重启问题解决
问题概述
我基于Python+FastAPI开发了多个交易机器人,每个机器人部署在独立的Google Cloud Run实例上,通过REST API控制。大部分时候运行正常,但偶尔会出现Uvicorn进程随机重启的情况,直接终止后台运行的循环交易函数,曾在交易中途发生重启,带来潜在风险。
相关代码与日志
Dockerfile
FROM python:3.10 COPY . /app WORKDIR /app RUN pip3 install -r requirements.txt CMD exec python3 -m uvicorn bot_handler:app --host 0.0.0.0 --port 8080
FastAPI启动接口
@app.get('/start',dependencies=[Depends(JWTBearer())]) def start(): thread=Thread(target=bot_instance.start) thread.daemon=True thread.start() return {"status":200}
重启时的系统日志
Info 2022-10-14 21:16:22.402 MSTGET200670 B3 mspython-requests/ Default 2022-10-14 21:42:26.421 MSTINFO: Shutting down Default 2022-10-14 21:42:26.523 MSTINFO: Waiting for application shutdown. Default 2022-10-14 21:42:26.523 MSTINFO: Application shutdown complete. Default 2022-10-14 21:42:26.523 MSTINFO: Finished server process [1] Default 2022-10-14 21:42:41.955 MSTINFO: Started server process [1] Default 2022-10-14 21:42:41.955 MSTINFO: Waiting for application startup. Default 2022-10-14 21:42:41.956 MSTINFO: Application startup complete. Default 2022-10-14 21:42:41.956 MSTINFO: Uvicorn running on http://0.0.0.0:8080 (Press CTRL+C to quit) Default 2022-10-14 22:58:52.745 MSTVALIDATING CONFIG FILE...
排查方向与解决方案
1. 确定Cloud Run重启的根本原因
Cloud Run实例自动重启通常有以下几种触发条件:
- 资源超限:CPU或内存使用率超过配置阈值,Cloud Run会强制重启实例。去Cloud Console的监控页面查看对应实例的资源使用曲线,若有超限情况,调高实例的CPU/内存配额。
- 自动缩容机制:默认情况下,Cloud Run在15分钟无请求后会将实例缩容到0,有新请求时再启动。但交易机器人是后台持续运行的任务,需要在Cloud Run服务设置里关闭“缩容至0实例”,把最小实例数设为1,确保实例一直运行。
- 健康检查失败:如果配置了存活探针或就绪探针,当应用无法在规定时间内响应检查请求,Cloud Run会重启实例。确保你的FastAPI应用提供了健康检查端点(比如
/health),并在Cloud Run设置中配置正确的检查路径和超时时间。
2. 优化交易任务的运行方式
- 替换守护线程:当前使用
daemon=True的线程运行交易逻辑,Uvicorn进程关闭时,守护线程会被直接杀死,没有机会完成当前交易或做收尾处理。改成非守护线程,并添加优雅退出逻辑:from fastapi import FastAPI, Depends import threading from contextlib import suppress app = FastAPI() bot_stop_flag = False bot_thread = None def run_bot(): global bot_stop_flag while not bot_stop_flag: # 执行交易循环逻辑 print("Executing trade cycle...") # 这里可以添加交易状态持久化代码,比如把当前步骤存到数据库 @app.get('/start', dependencies=[Depends(JWTBearer())]) def start(): global bot_thread, bot_stop_flag bot_stop_flag = False if bot_thread is None or not bot_thread.is_alive(): bot_thread = threading.Thread(target=run_bot) bot_thread.daemon = False bot_thread.start() return {"status":200} @app.on_event("shutdown") def handle_shutdown(): global bot_stop_flag, bot_thread bot_stop_flag = True if bot_thread is not None and bot_thread.is_alive(): # 等待线程完成当前交易循环,最多等30秒 with suppress(Exception): bot_thread.join(timeout=30) - 交易状态持久化:把交易的关键状态(比如当前订单ID、交易阶段、持仓信息等)存储到外部持久化存储(如Cloud Firestore、Cloud SQL),实例重启后,启动时先从存储中读取状态,恢复未完成的交易流程。
3. 调整Uvicorn启动参数
添加超时相关配置,避免因连接超时导致进程异常:
CMD exec python3 -m uvicorn bot_handler:app --host 0.0.0.0 --port 8080 --timeout-keep-alive 60 --graceful-timeout 30
内容的提问来源于stack exchange,提问作者Maanav
相关产品推荐
相关产品推荐

