You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Windows下Jupyter Notebook中CPU密集型ML推理的多进程加速问题

问题:16核CPU下Hugging Face模型推理提速不达预期

在Windows系统的VSCode Jupyter Notebook中使用Hugging Face模型处理一批句子,单句推理耗时约1秒,期望通过16核CPU实现高效加速。目前使用joblib仅将总推理时间从46.2秒降至29秒,提速效果远低于预期。尝试多种并行工具后,要么提速有限,要么出现报错、卡顿问题,相关代码及测试情况如下:

from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
import time

from joblib import Parallel, delayed
from tqdm.notebook import tqdm

# Input
sentences = [
    'The car is awful',
    'The party was amazing',
    'The dish is dirty',
    'My brother is the best',
    'My mom is the worst',
    'This example sucks',
    'The restaurant was so good',
    'I love the landscape',
    'Nature is wonderful',
    'John is terrible'
]

sentences = [x for y in [sentences] * 10 for x in y]

checkpoint = "kevinscaria/joint_tk-instruct-base-def-pos-neg-neut-combined"

tokenizer = AutoTokenizer.from_pretrained(checkpoint)
model = AutoModelForSeq2SeqLM.from_pretrained(checkpoint)

def absa(text):
    prompt = f"""Definition: The output will be the aspects (both implicit and explicit) and the aspects sentiment polarity.
Positive example 1-
input: I charge it at night and skip taking the cord with me because of the good battery life.
output: battery life:positive, 
Negative example 1-
input: Speaking of the browser, it too has problems.
output: browser:negative
Neutral example 1-
input: I took it back for an Asus and same thing- blue screen which required me to remove the battery to reset.
output: battery:neutral
Now complete the following example-
input: {text}\noutput: """
    tokenized_text = tokenizer(prompt, return_tensors="pt", truncation=True)
    output = model.generate(tokenized_text.input_ids, max_new_tokens=1000)
    detokenized = tokenizer.decode(output[0], skip_special_tokens=True)
    return detokenized

单线程测试(列表推导式)

%%time
[absa(x) for x in sentences]

# Wall time: 46.2 s

Joblib并行测试

%%time
Parallel(n_jobs=-1, backend='threading')(delayed(absa)(row) for row in tqdm(sentences))

# Wall time: 29 s

补充测试情况

对比含time.sleep()的虚拟函数与实际推理函数,测试多种并行工具的表现:

  • concurrent.futures.ThreadPoolExecutor:虚拟函数提速12倍,推理函数仅提速1.4倍
  • concurrent.futures.ProcessPoolExecutor:触发BrokenProcessPool错误
  • ray:虚拟函数提速11倍,推理函数触发"函数过大值"错误
  • multiprocess:运行时出现卡顿

内容的提问来源于stack exchange,提问作者zest16

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 01:02:03