You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用多进程生成Sentence Embedding时程序卡顿问题求助

问题分析与解决方案

你的代码卡顿无法执行,核心问题集中在以下几点:

1. 模型与Tokenizer跨进程传递失败

HuggingFace的模型和Tokenizer包含PyTorch张量、底层指针等无法被pickle序列化的结构,直接通过多进程参数传递会触发序列化异常或运行时死锁,这是卡顿的主要原因。

2. 异常捕获模块未导入

worker函数中使用except queue.Empty:但未导入queue模块,会触发NameError,导致worker进程崩溃或无限阻塞,无法处理后续任务。

3. 张量跨进程传递风险

直接将PyTorch张量存入mp.Manager().list(),可能因张量无法安全跨进程共享导致数据传递异常。


修复后的代码示例

让每个worker进程独立加载模型和Tokenizer,避免序列化问题,同时修复异常捕获和数据传递逻辑:

import torch
import multiprocessing as mp
from transformers import AutoTokenizer, AutoModel
from queue import Empty

def encode_sentence(sentence, model, tokenizer):
    encoded_input = tokenizer(sentence, return_tensors="pt")
    with torch.no_grad():
        output = model(**encoded_input)
    # 转为numpy数组,确保跨进程传递安全
    return output.last_hidden_state[:, 0, :].cpu().numpy()

def worker(input_queue, output_list, model_name):
    # 每个worker独立加载模型与Tokenizer
    tokenizer = AutoTokenizer.from_pretrained(model_name)
    model = AutoModel.from_pretrained(model_name)
    model.eval()
    
    while True:
        try:
            # 添加超时,避免队列空时无限阻塞
            sentence = input_queue.get(timeout=5)
            if sentence is None:
                break
            embedding = encode_sentence(sentence, model, tokenizer)
            output_list.append(embedding)
        except Empty:
            continue

if __name__ == "__main__":
    model_name = "distilbert-base-uncased"
    num_workers = mp.cpu_count()
    input_queue = mp.Queue()
    output_list = mp.Manager().list()
    
    # 传递模型名称而非已加载的模型对象
    workers = [mp.Process(target=worker, args=(input_queue, output_list, model_name)) for _ in range(num_workers)]
    for w in workers:
        w.start()

    sentences = ["This is sentence {}".format(i) for i in range(300)]
    for sentence in sentences:
        input_queue.put(sentence)

    # 给每个worker发送终止信号
    for _ in range(num_workers):
        input_queue.put(None)
    for w in workers:
        w.join()

    # 转回PyTorch张量(可选)
    embeddings = torch.tensor(list(output_list))
    print(embeddings.shape)

额外优化建议

如果你的环境使用GPU,多线程比多进程更适合(多进程会复制模型占用大量显存),可以改用ThreadPoolExecutor:

from concurrent.futures import ThreadPoolExecutor
import torch
from transformers import AutoTokenizer, AutoModel

def encode_sentence(sentence, model, tokenizer):
    encoded_input = tokenizer(sentence, return_tensors="pt")
    with torch.no_grad():
        output = model(**encoded_input)
    return output.last_hidden_state[:, 0, :]

if __name__ == "__main__":
    model_name = "distilbert-base-uncased"
    tokenizer = AutoTokenizer.from_pretrained(model_name)
    model = AutoModel.from_pretrained(model_name).eval()
    
    sentences = ["This is sentence {}".format(i) for i in range(300)]
    
    with ThreadPoolExecutor(max_workers=mp.cpu_count()) as executor:
        results = list(executor.map(lambda x: encode_sentence(x, model, tokenizer), sentences))
    
    embeddings = torch.cat(results, dim=0)
    print(embeddings.shape)

内容的提问来源于stack exchange,提问作者prashantgpt91

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 20:35:58