You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MacOS下运行LDA Mallet多进程触发freeze报错如何解决

问题根因

MacOS环境下Python 3.8及以上版本,多进程默认启动方式从fork切换为spawn。spawn模式启动子进程时会重新导入整个主模块,如果多进程调用代码写在模块顶层,子进程导入时会重复执行代码、反复拉起新进程,就会抛出当前看到的“启动阶段启动新进程”的RuntimeError。

修复方案

优先选第一种方案,兼容性最好:

  • 将所有业务执行逻辑、函数调用逻辑放到if __name__ == '__main__':判断块内。该判断块下的代码只会在主进程直接运行脚本时执行,子进程导入模块时不会触发,从根源避免重复拉起进程的问题。普通脚本运行不需要额外加freeze_support(),只有打包成exe等可执行文件时才需要加这行。

修正后的完整代码参考:

import gensim
from gensim.models.coherencemodel import CoherenceModel
from gensim.corpora import Dictionary
from gensim.models.ldamodel import LdaModel
import os.path

def optimize_parameters(lemma_tokens, texts):
    
    os.environ['MALLET_HOME'] = '****/mallet-2.0.8'
    mallet_path = '****/mallet-2.0.8/bin/mallet'

    id2word = Dictionary(lemma_tokens)

    # 过滤极端词
    id2word.filter_extremes(no_below=2, no_above=.99)

    # 生成语料对象
    corpus = [id2word.doc2bow(d) for d in lemma_tokens]

    model = gensim.models.wrappers.LdaMallet(mallet_path, corpus=corpus, num_topics=5, id2word=id2word, workers = 4)
    coherencemodel = CoherenceModel(model=model, texts=lemma_tokens, dictionary=id2word, coherence='c_v')
    coherence = coherencemodel.get_coherence()
    return coherence

if __name__ == '__main__':
    # 此处替换为你自己加载lemma_tokens、texts数据集的代码
    # lemma_tokens = 你的数据加载逻辑
    # texts = 你的数据加载逻辑
    coherence_score = optimize_parameters(lemma_tokens, texts)
  • 如果不想调整代码结构,可以在所有第三方库导入之前,手动将多进程启动方式切回fork,注意设置代码必须放在最开头,否则不生效:
import multiprocessing
# 必须在导入gensim等其他库之前执行
multiprocessing.set_start_method('fork', force=True)

# 后续原有导入、函数定义、执行逻辑保持不变即可

补充说明:部分新版MacOS的系统安全机制可能和fork模式存在兼容问题,偶发进程崩溃,优先使用第一种主入口判断的方案。如果调试阶段想快速验证流程,也可以先把LdaMallet初始化参数里的workers设为1,关闭多进程先跑通全链路,再调试多进程配置。

内容的提问来源于stack exchange,提问作者Yash

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 06:45:42