Spacy 3.2.0版本nlp.pipe多进程不生效,如何开启CPU并行处理?
问题根因
你遇到的多进程不生效问题,主要由以下几个常见原因导致:
- Windows系统运行时未将入口逻辑包裹在
if __name__ == "__main__"代码块中,spawn模式下无法正常创建子进程 - 你使用的
en_core_web_trf是Transformer类模型,依赖PyTorch运行时,模型本身不可序列化,无法跨进程传递,全局预加载后子进程无法拿到可用的模型实例,spacy会自动回退到单进程模式 - 传入的文本样本量过小,多进程启动开销高于处理收益,看不出耗时变化
- 你在
nlp.pipe中禁用了tok2vec等核心组件,实际参与运算的组件计算量极低,多核优势无法体现
解决方案
优先适配Windows运行环境
所有多进程相关的入口逻辑必须放在主程序判断块内,示例如下:import spacy import time if __name__ == "__main__": # 加载模型时直接禁用不需要的组件,效率高于在pipe中禁用 nlp = spacy.load("en_core_web_trf", disable=["tok2vec", "tagger", "parser", "attribute_ruler", "lemmatizer"]) docs = [你的文本列表] start_time = time.time() for doc in nlp.pipe(docs, n_process=16): # 业务操作 pass print("--- %s seconds --- " % (time.time() - start_time))Transformer模型优先用单进程多线程方案
Transformer模型本身支持PyTorch多线程调度,不需要开启spacy多进程就能跑满多核,代码修改如下:import spacy import time import torch # 设置PyTorch可用线程数,根据你的CPU核心数调整 torch.set_num_threads(16) nlp = spacy.load("en_core_web_trf", disable=["tok2vec", "tagger", "parser", "attribute_ruler", "lemmatizer"]) docs = [你的文本列表] start_time = time.time() # 不需要设置n_process参数 for doc in nlp.pipe(docs): # 业务操作 pass print("--- %s seconds --- " % (time.time() - start_time))必须使用多进程时的适配方案
如果一定要开启多进程处理,需要把模型加载逻辑放到子进程内执行,避免跨进程传递模型,示例如下:from multiprocessing import Pool def process_text(texts): import spacy nlp = spacy.load("en_core_web_trf", disable=["tok2vec", "tagger", "parser", "attribute_ruler", "lemmatizer"]) results = [] for doc in nlp.pipe(texts): # 处理逻辑,返回你需要的结果 results.append([ent.text for ent in doc.ents]) return results if __name__ == "__main__": docs = [你的文本列表] # 把文本切分成16份,对应16个进程 chunk_size = len(docs) // 16 chunks = [docs[i:i+chunk_size] for i in range(0, len(docs), chunk_size)] with Pool(16) as p: all_results = p.map(process_text, chunks)测试时保证样本量足够
测试多进程性能时,文本样本量至少要在1000条以上,避免进程启动开销掩盖处理收益。
内容的提问来源于stack exchange,提问作者Anton
相关产品推荐
相关产品推荐

