为何多核心训练Doc2Vec反而比单核心更慢?
多核心训练Doc2Vec耗时反而更长的问题
训练集train_corpus包含7930196条TaggedDocument格式的日志数据,示例如下:
print(train_corpus[:5]) [TaggedDocument(words=['port', 'ssh'], tags=[0]), TaggedDocument(words=['session', 'initialize', 'by', 'client'], tags=[1]), TaggedDocument(words=['dfs', 'fsnamesystem', 'block', 'namesystem', 'addstoredblock', 'blockmap', 'update', 'be', 'to', 'blk', 'size'], tags=[2]), TaggedDocument(words=['appl', 'selfupdate', 'component', 'amd', 'microsoft', 'windows', 'kernel', 'none', 'elevation', 'lower', 'version', 'revision', 'holder'], tags=[3]), TaggedDocument(words=['ramfs', 'tclass', 'blk', 'file'], tags=[4])]
环境信息
- 系统:CentOS7,8个可用核心
- Gensim版本:4.1.2
- BLAS库:OpenBLAS,已在bashrc中设置
OPENBLAS_NUM_THREADS=1,Jupyter中验证配置生效 - 系统无其他负载,仅运行训练任务
测试代码
dict_time_workers = dict() for workers in range(1, 9): model = Doc2Vec(vector_size=20, min_count=1, workers=workers, epochs=1) model.build_vocab(train_corpus, update = False) t1 = time.time() model.train(train_corpus, epochs=1, total_examples=model.corpus_count) dict_time_workers[workers] = time.time() - t1
训练耗时结果
{1: 224.23211407661438, 2: 273.408652305603, 3: 313.1667754650116, 4: 331.1840877532959, 5: 433.83785605430603, 6: 545.671571969986, 7: 551.6248495578766, 8: 548.430994272232}
异常现象
- 核心数越多,训练耗时越长,增加
epochs参数后结果一致 - 通过htop观察,线程数与设置的核心数对应,但核心使用率随核心数增加而降低:单核心使用率95%,双核心各65%,六核心各20-25%
- iotop检测磁盘无异常,排除IO瓶颈
内容的提问来源于stack exchange,提问作者Naindlac
相关产品推荐
相关产品推荐

