You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让多Gunicorn进程共享已加载NLP模型的字典?

问题背景与现状
  • 运行NLP模型的Flask服务,连接设备从5台增至15台后出现性能瓶颈
  • 服务器核心数从12增至24,性能仅提升约20%
  • 改用Gunicorn+taskset启动3个进程(每个分配8核、限5台设备),性能回到原12核5设备水平,但内存消耗增至3倍——原因是每个进程独立加载相同模型

当前代码实现

模型加载代码

import torch
import torch.nn as nn
import torchvision
import os
import learning

app = Flask(__name__)

models_list_name = os.listdir('models')

global_dict_models = {}

for token in models_list_name:
    try:
        tok, n_class, model_type, typ, h_layer = token.split('.')

        t, model = learning.get_model(model_type, typ)

        n_class = int(n_class)
        h_layer = int(h_layer)
        model = learning.NewModel(model, n_class, typ, h_layer)
        model.load_state_dict(torch.load(os.path.join('models', token), map_location=torch.device('cpu')))
        
        print(model, tok)
        with open(tok, 'r') as f:
            label = f.read().split('\n')
        print(label)
        global_dict_models[tok] = (t, model, label)
    except ValueError:
        pass          

预测调用代码

t, model, labels = global_dict_models[token]
x = t.encode_plus(text, add_special_tokens=True, max_length=512, truncation=True, padding="max_length", return_tensors='pt')
    
output = torch.sigmoid(model(x['input_ids'].squeeze(1), x['attention_mask'])).detach().cpu().numpy()

封装后的尝试代码(存在进程共享问题)

from flask import Flask, request, make_response
import torch
import torch.nn as nn
import torchvision
import os
import learning

app = Flask(__name__)

models_list_name = os.listdir('models')

global_dict_models = {}

def load_models():
    for token in models_list_name:
        try:
            tok, n_class, model_type, typ, h_layer = token.split('.')
    
            t, model = learning.get_model(model_type, typ)
    
            n_class = int(n_class)
            h_layer = int(h_layer)
            model = learning.NewModel(model, n_class, typ, h_layer)
            model.load_state_dict(torch.load(os.path.join('models', token), map_location=torch.device('cpu')))
            
            print(model, tok)
            with open(tok, 'r') as f:
                label = f.read().split('\n')
            print(label)
            global_dict_models[tok] = (t, model, label)
        except ValueError:
            pass
  

def dict_models(token):
    t, model, labels = global_dict_models[token]
    return t, model, labels    
    
if __name__ == '__main__':
    load_models()
    print('Start')
    app.run(host="0.0.0.0", port="5000", threaded=True, processes=1)

核心问题:封装后多进程无法共享已加载的模型字典,需要实现模型一次加载、多进程共享,或其他性能优化方案

解决方案

1. 利用PyTorch共享内存机制(CPU模型)

PyTorch支持将模型张量放入共享内存,子进程可直接访问无需重复加载:

  • 加载模型时调用model.share_memory(),将所有参数转为共享内存张量
  • 使用Gunicorn时添加--preload参数,让主进程先加载模型,子进程继承共享内存中的模型实例
  • 注意:分词器t属于不可序列化对象,需每个子进程单独初始化,或提前将分词器词汇表存入共享存储(如本地文件)

修改后的加载逻辑示例:

def load_models():
    for token in models_list_name:
        try:
            tok, n_class, model_type, typ, h_layer = token.split('.')
    
            t, model = learning.get_model(model_type, typ)
            n_class = int(n_class)
            h_layer = int(h_layer)
            model = learning.NewModel(model, n_class, typ, h_layer)
            model.load_state_dict(torch.load(os.path.join('models', token), map_location=torch.device('cpu')))
            # 转为共享内存
            model.share_memory()
            
            print(model, tok)
            with open(tok, 'r') as f:
                label = f.read().split('\n')
            # 仅存储模型和标签,分词器由子进程自行初始化
            global_dict_models[tok] = (model, label, model_type, typ)
        except ValueError:
            pass

启动Gunicorn命令:

gunicorn --preload --workers 3 --bind 0.0.0.0:5000 your_app:app

子进程预测时初始化分词器:

# 预测逻辑
model, labels, model_type, typ = global_dict_models[token]
# 子进程单独初始化分词器
t, _ = learning.get_model(model_type, typ)
x = t.encode_plus(text, add_special_tokens=True, max_length=512, truncation=True, padding="max_length", return_tensors='pt')
output = torch.sigmoid(model(x['input_ids'].squeeze(1), x['attention_mask'])).detach().cpu().numpy()

2. 采用模型服务化架构(如TorchServe)

将NLP模型独立部署为模型服务,Flask仅作为请求网关:

  • 使用TorchServe打包模型,启动单独的模型服务进程,该服务已优化并发处理和资源利用
  • Flask负责接收设备请求,将文本转发给TorchServe获取预测结果,再返回给设备
  • 优势:模型仅加载一次,无需手动处理进程共享问题,扩展性更强

3. 切换为多线程模式(规避多进程重复加载)

CPU运行的PyTorch模型会自动释放GIL,多线程可有效利用多核资源:

  • 用Gunicorn启动单进程多线程服务,设置--threads参数(如--threads 12)
  • 模型仅加载一次,所有线程共享同一模型字典,内存消耗低
  • 注意:确保预测代码线程安全,PyTorch推理默认线程安全(仅读取模型参数时)

启动命令示例:

gunicorn --workers 1 --threads 12 --bind 0.0.0.0:5000 your_app:app

4. 基于multiprocessing.Manager的共享字典(复杂方案)

通过Python的multiprocessing.Manager创建共享字典,但需结合共享内存张量:

  • 主进程加载模型后,将模型参数转为共享内存张量,把张量引用存入共享字典
  • 子进程从字典获取张量引用,重新组装模型
  • 缺点:实现复杂,不如PyTorch原生share_memory()方案直接

内容的提问来源于stack exchange,提问作者Роман Чирков

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 15:06:29