FastAPI集成HuggingFace Transformer的子进程与单Worker崩溃问题
问题与解答
问题背景
使用FastAPI开发基于HuggingFace RobertaForQuestionAnswering模型的问答API时,遇到两个关键问题:
- 每次调用模型接口都会启动新子进程并重新加载模型,耗时极长,尝试过依赖注入、单例、lifespan等方案均无效。
- 单Worker运行uvicorn时,调用模型接口会触发API优雅关闭,仅设置
workers=2能临时缓解。
相关代码片段:
tokenizer = RobertaTokenizer.from_pretrained(data_dir) model = RobertaForQuestionAnswering.from_pretrained(data_dir) transformer_pipeline = pipeline(task="question-answering", model=model, tokenizer=tokenizer) @router.post("/test") async def test(): return response_model(question="What is my name?", context="My name is Joe")
单Worker下调用模型后的控制台输出:
######### Start the app INFO: Application startup complete. INFO: Uvicorn running on http://0.0.0.0:8080 (Press CTRL+C to quit) ######### 调用非模型接口正常 ... ######### 调用/test2接口(返回结果但应用关闭) INFO: Shutting down INFO: Waiting for connections to close. (CTRL+C to force quit) INFO: 127.0.0.1:39138 - "POST /test2 HTTP/1.1" 200 OK INFO: Waiting for application shutdown. in asynccontextmanager shutdown block INFO: Application shutdown complete. INFO: Finished server process [126197]
已尝试方案:Depends依赖注入、等待模型加载、@asynccontextmanager、@app.on_event("startup")预加载、单例类。
核心疑问解答
1. 为何调用模型时会创建子进程?
FastAPI本身不会主动创建子进程,问题根源是HuggingFace Transformers的pipeline默认行为:
pipeline初始化时若未明确指定设备或多进程参数,会自动检测环境。当检测到CUDA可用时,部分场景会启动子进程处理推理;即使CPU环境,部分pipeline实现也会通过子进程隔离推理逻辑。- 若开启了uvicorn的
--reload热重载,watchgod监控文件变化会触发进程重启,但你的情况禁用重载后问题依旧,所以核心是pipeline的多进程机制。
2. 单Worker下调用模型为何触发应用关闭?
单Worker模式下,uvicorn主进程同时负责请求接收与处理,Transformers子进程的信号会与uvicorn的信号处理逻辑冲突:
- 子进程退出时可能发送SIGINT/SIGTERM信号,被uvicorn主进程捕获,误以为是用户触发的关闭指令,从而启动优雅关闭流程。
- 异步接口(
async def)会让FastAPI用线程池处理同步的模型推理,进一步加剧信号冲突的概率。
解决方案
1. 强制pipeline使用单进程模式
初始化pipeline时明确禁用多进程,指定设备:
transformer_pipeline = pipeline( task="question-answering", model=model, tokenizer=tokenizer, device=-1, # 强制用CPU;用GPU则设为0 use_multiprocessing=False, # 禁用多进程 num_workers=0 # 确保不启动 worker 进程 )
2. 改用同步接口处理推理
Transformers的pipeline是同步操作,将接口改为def而非async def,避免线程池带来的潜在冲突:
@router.post("/test") def test(): result = transformer_pipeline(question="What is my name?", context="My name is Joe") return result
3. 规范使用Lifespan管理模型状态
确保模型在应用启动时仅加载一次,关闭时正确清理:
from contextlib import asynccontextmanager from fastapi import FastAPI @asynccontextmanager async def lifespan(app: FastAPI): # 启动加载模型 global tokenizer, model, transformer_pipeline tokenizer = RobertaTokenizer.from_pretrained(data_dir) model = RobertaForQuestionAnswering.from_pretrained(data_dir) transformer_pipeline = pipeline( task="question-answering", model=model, tokenizer=tokenizer, use_multiprocessing=False ) yield # 关闭清理资源 del model, tokenizer, transformer_pipeline app = FastAPI(lifespan=lifespan)
4. 临时禁用uvicorn信号处理
若以上方法无效,启动uvicorn时添加参数避免信号误捕获:
uvicorn main:app --host 0.0.0.0 --port 8080 --no-sigint-handler --no-sigterm-handler
内容的提问来源于stack exchange,提问作者banjoeschmoe
相关产品推荐
相关产品推荐

