Docker容器中终止进程致Uvicorn崩溃的问题与优化咨询
问题描述
我开发了一个数据处理Worker服务,核心逻辑如下:
- 发送POST请求到指定路由时,启动主进程,主进程内为每个数据表开启子进程执行并行循环任务
- 发送DELETE请求到指定路由时,终止指定进程
本地环境中,终止主进程时子进程会同步终止,但在Docker容器中执行终止操作后出现两个异常:
- Uvicorn服务器自行关闭,服务无法接收新请求
- 目标主进程的子进程仍在后台残留运行
目前尝试集成Tini并使用SIGKILL替代SIGTERM可以终止进程,但不确定该方案的合理性,寻求优化建议。
启动路由代码
@worker.post('/process/production-process-monthly/start', tags=["worker"]) async def calculate_solar_monthly( config: Settings = Depends(get_setting), body: Union[CalculateSolarEnergyMonthly, None] = Body( None, title="These values are for choosing the calculation model of the accumulated data.", description="Refers to units not like electricity, natural gas, oil.", alias="body"), # token: dict = Depends(token_decoder)): try: log.info(f"-> Calculating operation time") log.warning("-" * 60) start = time.time() business = 'hktm' # we will get this from token. # return an error if there is a process with the same ID as a process in the processes list process_id = f"{config.keyspace_name}.{business}_monthly_test" existing_process = next( (p for p in processes if p["id"] == process_id), None) if existing_process: return JSONResponse(content={"message": f"{process_id} is not RUNNING. Duplicated error."}, status_code=500) # run main function in parallel using multiprocessing.Process p = multiprocessing.Process( target=monthlyProcess, args=(config, body, business)) p.start() # run main function in parallel using multiprocessing.Process processes.append({"p": p, "id": process_id}) except: return JSONResponse( content=config.err.WORKER_PROCESS_START_ERROR, status_code=500, )
子进程循环代码
for table in table_names: if table.endswith('daily_testx'): tables.append(table) # List to store subprocesses # newTable = [] # newTable.append(tables[0]) while True: start_time = time.time() with multiprocessing.Pool(processes=len(tables)) as pool: pool.starmap(startMonthlyProcess, [ (table, config, body, business) for table in tables]) pool.close() pool.join() log.warning( "Process finished --- %s seconds ---" % (time.time() - start_time)) time.sleep(body.period * 60)
停止路由代码
@worker.delete('/process/production-process-monthly/stop', tags=["worker"]) async def calculate_solar_daily( config: Settings = Depends(get_setting), process_name: Union[str, None] = Query( None, title="Update key", description="Update object values.", alias="process_name"), key_prefix: Union[str, None] = Query( None, title="Update key", description="Update object values.", alias="key_prefix")): if process_name == None: return JSONResponse(content={"message": f"{process_name} is not Stopped. The variable contains a NULL value."}, status_code=500) try: delete_redis_keys_of_pa(config, key_prefix) for idx, p in enumerate(processes): if p["id"].endswith(process_name): # print(p["p"]) p = p["p"] log.info(p) # print('p in,', p) # parentpid = p._parent_pid try: os.kill(p.pid, signal.SIGTERM) except Exception as ex: log.info(f"### {ex} happened") # p.terminate() processes.pop(idx) log.info( colored("message: Stopped process with ID ", "blue") + colored(f"{p.pid}", "yellow") + colored(" Stopped", "red"))
异常日志
2023-09-25 13:26:20 0bcd0b1e3552 Worker ROUTE[1] INFO <Process name='Process-1' pid=27 parent=1 started> 2023-09-25 13:26:20 0bcd0b1e3552 Worker ROUTE[1] INFO message: Stopped process with ID 27 Stopped 2023-09-25 13:26:20 0bcd0b1e3552 uvicorn.error[1] INFO Shutting down INFO: 192.168.128.11:42780 - "DELETE /api/v1/cost/process/production-process-monthly/stop?process_name=monthly_test&key_prefix=daily_testx HTTP/1.0" 200 OK 2023-09-25 13:26:21 0bcd0b1e3552 uvicorn.error[1] INFO Waiting for application shutdown. 2023-09-25 13:26:21 0bcd0b1e3552 uvicorn.error[1] INFO Application shutdown complete. 2023-09-25 13:26:21 0bcd0b1e3552 uvicorn.error[1] INFO Finished server process [1]
分析与优化建议
问题根源分析
- Uvicorn退出原因:Docker容器中,Uvicorn作为PID 1进程运行时,信号处理机制与本地环境不同。PID 1进程默认不会转发未处理的信号,且当子进程终止时,若信号处理不当可能触发Uvicorn自身的退出逻辑。从日志看,主进程PID 27的父进程是1(Uvicorn),终止主进程时可能导致信号传播到PID 1,触发Uvicorn关闭。
- 子进程残留原因:主进程被终止后,其下属的子进程成为孤儿进程,Docker中若PID 1是Uvicorn,它不会主动清理这些孤儿进程,导致子进程后台残留运行。
优化方案
1. 进程组管理:一次性终止主进程及所有子进程
启动主进程时,将其设为新进程组的组长,终止时发送信号给整个进程组,确保所有子进程同步终止:
- 修改启动代码,添加
start_new_session=True:
p = multiprocessing.Process( target=monthlyProcess, args=(config, body, business), start_new_session=True # 让主进程成为新进程组组长 )
- 修改停止代码,发送信号到进程组而非单个进程:
# 替换原os.kill(p.pid, signal.SIGTERM) os.killpg(os.getpgid(p.pid), signal.SIGTERM)
2. 用Tini作为Docker容器的PID 1进程
在Docker容器中使用Tini作为init进程,负责信号转发和孤儿进程清理,避免Uvicorn直接作为PID 1:
# Dockerfile中添加Tini RUN apt-get update && apt-get install -y tini ENTRYPOINT ["/usr/bin/tini", "--"] CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"]
3. 子进程优雅终止逻辑
在子进程中添加信号处理函数,捕获SIGTERM信号后优雅关闭进程池,避免任务中断导致的数据异常:
# 在monthlyProcess函数开头添加 import signal import sys def sigterm_handler(signum, frame): log.info("Received SIGTERM, starting graceful shutdown...") # 若使用全局进程池引用,需提前定义pool变量 global pool if 'pool' in globals() and pool is not None: pool.terminate() sys.exit(0) signal.signal(signal.SIGTERM, sigterm_handler)
4. 避免滥用SIGKILL
SIGKILL会强制终止进程,可能导致数据丢失或资源泄漏,仅当SIGTERM无法终止进程时再作为兜底方案使用。优先通过进程组管理+优雅终止逻辑实现正常关闭。
5. 进程列表维护优化
定期清理进程列表中已终止的进程,避免残留无效进程对象:
# 在启动新进程前,过滤掉已终止的进程 processes = [proc for proc in processes if proc["p"].is_alive()]
内容的提问来源于stack exchange,提问作者Furkan YIlmaZ
相关产品推荐
相关产品推荐

