You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何多核心训练Doc2Vec反而比单核心更慢?

多核心训练Doc2Vec耗时反而更长的问题

训练集train_corpus包含7930196条TaggedDocument格式的日志数据,示例如下:

print(train_corpus[:5])

[TaggedDocument(words=['port', 'ssh'], tags=[0]),
 TaggedDocument(words=['session', 'initialize', 'by', 'client'], tags=[1]),
 TaggedDocument(words=['dfs', 'fsnamesystem', 'block', 'namesystem', 'addstoredblock', 'blockmap', 'update', 'be', 'to', 'blk', 'size'], tags=[2]),
 TaggedDocument(words=['appl', 'selfupdate', 'component', 'amd', 'microsoft', 'windows', 'kernel', 'none', 'elevation', 'lower', 'version', 'revision', 'holder'], tags=[3]),
 TaggedDocument(words=['ramfs', 'tclass', 'blk', 'file'], tags=[4])]

环境信息

  • 系统:CentOS7,8个可用核心
  • Gensim版本:4.1.2
  • BLAS库:OpenBLAS,已在bashrc中设置OPENBLAS_NUM_THREADS=1,Jupyter中验证配置生效
  • 系统无其他负载,仅运行训练任务

测试代码

dict_time_workers = dict()
for workers in range(1, 9):

    model =  Doc2Vec(vector_size=20,
                    min_count=1,
                    workers=workers,
                    epochs=1)
    model.build_vocab(train_corpus, update = False)
    t1 = time.time()
    model.train(train_corpus, epochs=1, total_examples=model.corpus_count) 
    dict_time_workers[workers] = time.time() - t1

训练耗时结果

{1: 224.23211407661438, 
2: 273.408652305603, 
3: 313.1667754650116, 
4: 331.1840877532959, 
5: 433.83785605430603,
6: 545.671571969986, 
7: 551.6248495578766, 
8: 548.430994272232}

异常现象

  • 核心数越多,训练耗时越长,增加epochs参数后结果一致
  • 通过htop观察,线程数与设置的核心数对应,但核心使用率随核心数增加而降低:单核心使用率95%,双核心各65%,六核心各20-25%
  • iotop检测磁盘无异常,排除IO瓶颈

内容的提问来源于stack exchange,提问作者Naindlac

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 16:31:00