You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

FastAPI部署GPU型ML模型API:并行请求与CUDA内存问题求助

解决GPU模型API的并行请求与CUDA内存问题

一、修正代码基础错误

你的Model_loading类构造方法缺少self参数,会导致实例化失败,先修正:

class Model_loading():
    def __init__(self):  # 添加self参数
        self.model = torch.load('model.pth')
        self.model.eval()  # 必须设置为eval模式,关闭梯度计算

二、实现并行请求高效处理

当前3个请求串行执行9秒,核心原因是GPU任务没有真正并行,结合FastAPI和PyTorch的特性,按以下步骤调整:

1. 启用Uvicorn多进程Worker

FastAPI同步端点(def定义)默认在单进程线程池处理,但GPU计算需要多进程来充分利用并行能力。启动Uvicorn时指定多worker:

uvicorn main:app --workers 2 --host 0.0.0.0 --port 8000

注意:worker数量不要超过GPU能承载的模型副本数(每个worker会加载一份模型到GPU),比如16GB显存的GPU,若单模型占4GB,最多开4个worker。

2. 用线程池封装GPU计算

PyTorch的CUDA操作会自动释放GIL,线程池可实现GPU任务并行。用fastapi.concurrency.run_in_threadpool封装GPU计算,避免阻塞主线程:

from fastapi.concurrency import run_in_threadpool

@app.post('/model-testing')
async def my_function(file: UploadFile = File(...)):
    # 预处理(CPU操作,可同步执行)
    x = ...  # 你的预处理逻辑
    
    # 用线程池执行GPU计算,释放主线程
    output = await run_in_threadpool(gpu_based_processing, x)
    
    return Response(content=output , media_type="image/jpg")

同步修改GPU处理函数,强制禁用梯度计算:

def gpu_based_processing(x):
    with torch.no_grad():  # 禁用梯度,减少内存占用并加速
        # GPU计算逻辑
        result = model_instance.model(x.to('cuda'))
        # 将结果转回CPU,避免显存占用
        return result.cpu().numpy()  # 或转换为你需要的输出格式

3. 多进程模型内存共享(可选)

如果多worker导致模型重复加载占用显存,用torch.multiprocessing实现模型内存共享:

import torch.multiprocessing as mp

def load_model():
    model = torch.load('model.pth').eval()
    return model

if __name__ == "__main__":
    mp.set_start_method('spawn')
    model = load_model()
    model.share_memory()  # 共享模型内存
    model_instance = Model_loading(model)  # 修改类接收外部模型
    uvicorn.run("main:app", workers=2)

对应的Model_loading类调整:

class Model_loading():
    def __init__(self, model):
        self.model = model

三、解决CUDA内存不足问题

1. 即时释放显存

每个请求处理完后,手动删除无用张量并清空缓存:

def gpu_based_processing(x):
    with torch.no_grad():
        x_tensor = x.to('cuda')
        result = model_instance.model(x_tensor)
        # 立即释放输入张量
        del x_tensor
        torch.cuda.empty_cache()  # 清空无用显存
        return result.cpu().numpy()

2. 模型半精度量化

将模型转为FP16半精度,显存占用直接减半:

class Model_loading():
    def __init__(self):
        self.model = torch.load('model.pth').eval().half()  # 转半精度

预处理时同步将输入转为半精度:

x_tensor = x.to('cuda').half()

3. 限制并发请求数

用Semaphore控制同时处理的请求数,避免超过GPU内存上限:

from asyncio import Semaphore

# 根据GPU显存设置最大并发数,比如4
sem = Semaphore(4)

@app.post('/model-testing')
async def my_function(file: UploadFile = File(...)):
    async with sem:
        # 预处理逻辑
        x = ...
        output = await run_in_threadpool(gpu_based_processing, x)
        return Response(content=output , media_type="image/jpg")

4. 及时清理上传文件

处理完上传文件后立即关闭,释放内存:

async def my_function(file: UploadFile = File(...)):
    async with sem:
        contents = await file.read()
        await file.close()  # 及时关闭文件释放内存
        # 预处理逻辑
        x = ...

四、验证效果

调整后,3个并行请求应能在3秒左右完成;20个请求时通过Semaphore限制并发数,可避免CUDA内存不足错误。

内容的提问来源于stack exchange,提问作者dummyuser-123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 01:35:32