You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

FastAPI集成HuggingFace Transformer的子进程与单Worker崩溃问题

问题与解答

问题背景

使用FastAPI开发基于HuggingFace RobertaForQuestionAnswering模型的问答API时,遇到两个关键问题:

  1. 每次调用模型接口都会启动新子进程并重新加载模型,耗时极长,尝试过依赖注入、单例、lifespan等方案均无效。
  2. 单Worker运行uvicorn时,调用模型接口会触发API优雅关闭,仅设置workers=2能临时缓解。

相关代码片段:

tokenizer = RobertaTokenizer.from_pretrained(data_dir)
model = RobertaForQuestionAnswering.from_pretrained(data_dir)
transformer_pipeline = pipeline(task="question-answering", model=model, tokenizer=tokenizer)

@router.post("/test")
async def test():
    return response_model(question="What is my name?", context="My name is Joe")

单Worker下调用模型后的控制台输出:

######### Start the app
INFO:     Application startup complete.
INFO:     Uvicorn running on http://0.0.0.0:8080 (Press CTRL+C to quit)
######### 调用非模型接口正常
...
######### 调用/test2接口(返回结果但应用关闭)
INFO:     Shutting down
INFO:     Waiting for connections to close. (CTRL+C to force quit)
INFO:     127.0.0.1:39138 - "POST /test2 HTTP/1.1" 200 OK
INFO:     Waiting for application shutdown.
in asynccontextmanager shutdown block
INFO:     Application shutdown complete.
INFO:     Finished server process [126197]

已尝试方案:Depends依赖注入、等待模型加载、@asynccontextmanager、@app.on_event("startup")预加载、单例类。


核心疑问解答

1. 为何调用模型时会创建子进程?

FastAPI本身不会主动创建子进程,问题根源是HuggingFace Transformers的pipeline默认行为:

  • pipeline初始化时若未明确指定设备或多进程参数,会自动检测环境。当检测到CUDA可用时,部分场景会启动子进程处理推理;即使CPU环境,部分pipeline实现也会通过子进程隔离推理逻辑。
  • 若开启了uvicorn的--reload热重载,watchgod监控文件变化会触发进程重启,但你的情况禁用重载后问题依旧,所以核心是pipeline的多进程机制。

2. 单Worker下调用模型为何触发应用关闭?

单Worker模式下,uvicorn主进程同时负责请求接收与处理,Transformers子进程的信号会与uvicorn的信号处理逻辑冲突:

  • 子进程退出时可能发送SIGINT/SIGTERM信号,被uvicorn主进程捕获,误以为是用户触发的关闭指令,从而启动优雅关闭流程。
  • 异步接口(async def)会让FastAPI用线程池处理同步的模型推理,进一步加剧信号冲突的概率。

解决方案

1. 强制pipeline使用单进程模式

初始化pipeline时明确禁用多进程,指定设备:

transformer_pipeline = pipeline(
    task="question-answering",
    model=model,
    tokenizer=tokenizer,
    device=-1,  # 强制用CPU;用GPU则设为0
    use_multiprocessing=False,  # 禁用多进程
    num_workers=0  # 确保不启动 worker 进程
)

2. 改用同步接口处理推理

Transformers的pipeline是同步操作,将接口改为def而非async def,避免线程池带来的潜在冲突:

@router.post("/test")
def test():
    result = transformer_pipeline(question="What is my name?", context="My name is Joe")
    return result

3. 规范使用Lifespan管理模型状态

确保模型在应用启动时仅加载一次,关闭时正确清理:

from contextlib import asynccontextmanager
from fastapi import FastAPI

@asynccontextmanager
async def lifespan(app: FastAPI):
    # 启动加载模型
    global tokenizer, model, transformer_pipeline
    tokenizer = RobertaTokenizer.from_pretrained(data_dir)
    model = RobertaForQuestionAnswering.from_pretrained(data_dir)
    transformer_pipeline = pipeline(
        task="question-answering",
        model=model,
        tokenizer=tokenizer,
        use_multiprocessing=False
    )
    yield
    # 关闭清理资源
    del model, tokenizer, transformer_pipeline

app = FastAPI(lifespan=lifespan)

4. 临时禁用uvicorn信号处理

若以上方法无效,启动uvicorn时添加参数避免信号误捕获:

uvicorn main:app --host 0.0.0.0 --port 8080 --no-sigint-handler --no-sigterm-handler

内容的提问来源于stack exchange,提问作者banjoeschmoe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 08:10:43