You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Scikit-learn中的LDA主题模型结果可复现?

如何让scikit-learn的LatentDirichletAllocation模型结果可复现?

我正在用LDA进行主题建模,通过from sklearn.decomposition import LatentDirichletAllocation导入模块,基于10份文件构建模型并尝试聚类为3个主题。但每次运行时,即使输入数据完全相同,聚类结果也有差异,模型不具备可复现性。

我的示例代码如下:

import numpy as np 
data = [] 
a1 = " a word in groupa doca" 
a2 = " a word in groupa docb" 
a3 = "a word in groupb docc" 
a4 = "a word in groupc docd" 
a5 ="a word in groupc doce" 
data = [a1,a2,a3,a4,a5] 
del a1,a2,a3,a4,a5 
NO_DOCUMENTS = len(data) 
print(NO_DOCUMENTS) 

from sklearn.decomposition import LatentDirichletAllocation 
from sklearn.feature_extraction.text import CountVectorizer 

NUM_TOPICS = 2 
vectorizer = CountVectorizer(min_df=0.001, max_df=0.99998, stop_words='english', lowercase=True, token_pattern='[a-zA-Z\-][a-zA-Z\-]{2,}') 
data_vectorized = vectorizer.fit_transform(data) 

# Build a Latent Dirichlet Allocation Model 
lda_model = LatentDirichletAllocation(n_topics=NUM_TOPICS, max_iter=10, learning_method='online') 
lda_Z = lda_model.fit_transform(data_vectorized) 

vocab = vectorizer.get_feature_names() 
text = "The economy is working better than ever" 
x = lda_model.transform(vectorizer.transform([text]))[0] 
print(x, x.sum()) 

TOPICWISEDOCUMENTS = {}
for iDocIndex,text in enumerate(data): 
    x = list(lda_model.transform(vectorizer.transform([text]))[0]) 
    maxIndex = x.index(max(x)) 
    if maxIndex in TOPICWISEDOCUMENTS: 
        TOPICWISEDOCUMENTS[maxIndex].append(iDocIndex) 
    else: 
        TOPICWISEDOCUMENTS[maxIndex] = [iDocIndex] 
print(TOPICWISEDOCUMENTS)

这个问题很常见,原因是LatentDirichletAllocation模型在初始化和训练过程中引入了随机操作(比如主题分布的随机初始化、在线学习时的随机抽样等),所以每次运行的结果会有差异。要实现可复现,只需要固定所有涉及随机数生成的环节即可,具体步骤如下:

1. 设置全局随机种子

在代码最开始的地方,设置numpy的随机种子,因为scikit-learn的很多随机操作依赖于numpy的随机数生成器:

import numpy as np
np.random.seed(42)  # 42是常用的种子值,你可以换成任意整数

2. 给LDA模型指定random_state参数

在初始化LatentDirichletAllocation的时候,添加random_state参数,固定模型内部的随机状态:

lda_model = LatentDirichletAllocation(
    n_topics=NUM_TOPICS, 
    max_iter=10, 
    learning_method='online',
    random_state=42  # 和全局种子一致或者用其他固定整数都可以
)

3. 额外注意(如果使用多线程)

如果你开启了多线程训练(比如设置了n_jobs>1),还需要设置sklearn的全局随机种子,避免多线程环境下的随机差异:

import sklearn
sklearn.random.seed(42)

修正后的完整可复现代码

把这些修改整合到你的代码里,最终版本如下:

import numpy as np 
np.random.seed(42)  # 设置全局numpy随机种子
import sklearn
sklearn.random.seed(42)  # 针对多线程场景的额外保障

data = [] 
a1 = " a word in groupa doca" 
a2 = " a word in groupa docb" 
a3 = "a word in groupb docc" 
a4 = "a word in groupc docd" 
a5 ="a word in groupc doce" 
data = [a1,a2,a3,a4,a5] 
del a1,a2,a3,a4,a5 
NO_DOCUMENTS = len(data) 
print(NO_DOCUMENTS) 

from sklearn.decomposition import LatentDirichletAllocation 
from sklearn.feature_extraction.text import CountVectorizer 

NUM_TOPICS = 2 
vectorizer = CountVectorizer(min_df=0.001, max_df=0.99998, stop_words='english', lowercase=True, token_pattern='[a-zA-Z\-][a-zA-Z\-]{2,}') 
data_vectorized = vectorizer.fit_transform(data) 

# Build a Latent Dirichlet Allocation Model with fixed random state
lda_model = LatentDirichletAllocation(
    n_topics=NUM_TOPICS, 
    max_iter=10, 
    learning_method='online',
    random_state=42
) 
lda_Z = lda_model.fit_transform(data_vectorized) 

vocab = vectorizer.get_feature_names() 
text = "The economy is working better than ever" 
x = lda_model.transform(vectorizer.transform([text]))[0] 
print(x, x.sum()) 

TOPICWISEDOCUMENTS = {}
for iDocIndex,text in enumerate(data): 
    x = list(lda_model.transform(vectorizer.transform([text]))[0]) 
    maxIndex = x.index(max(x)) 
    if maxIndex in TOPICWISEDOCUMENTS: 
        TOPICWISEDOCUMENTS[maxIndex].append(iDocIndex) 
    else: 
        TOPICWISEDOCUMENTS[maxIndex] = [iDocIndex] 
print(TOPICWISEDOCUMENTS)

这样修改之后,每次运行代码都会得到完全一致的聚类结果和主题分布啦。


内容的提问来源于stack exchange,提问作者rajeshkumargp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:24:15