如何计算SkLearn中LSA与LDA模型的Coherence Score
scikit-learn本身未内置主题一致性(Coherence Score)计算能力,目前业内通用方案是借助gensim库的CoherenceModel实现计算,具体操作步骤如下:
前置数据准备
计算Coherence Score需要先准备三类数据:原始分词后的文本列表、词表词典、词袋格式语料,你可以基于现有向量化逻辑直接转换:
from gensim.corpora.dictionary import Dictionary from gensim.models import CoherenceModel # 你做文本向量化用的Vectorizer,假设存储在vectorizer变量中 feature_names = vectorizer.get_feature_names_out() # 原始分词后的文本列表,每个元素为单条文本的分词结果列表,假设存储在texts变量中 dictionary = Dictionary(texts) corpus = [dictionary.doc2bow(text) for text in texts] # 统一设置每个主题取前20个关键词计算一致性,可根据需求调整 TOP_N_KEYWORDS = 20
提取两类模型的主题关键词列表
LSA模型主题提取
lsa_topics = [] for topic in lsa_model.components_: # 取权重最高的TOP_N_KEYWORDS个词作为当前主题的代表词 top_idx = topic.argsort()[-TOP_N_KEYWORDS:] lsa_topics.append([feature_names[i] for i in top_idx])
LDA模型主题提取
scikit-learn中LDA和LSA的components_属性结构完全一致,提取逻辑相同:
lda_topics = [] for topic in lda_model.components_: top_idx = topic.argsort()[-TOP_N_KEYWORDS:] lda_topics.append([feature_names[i] for i in top_idx])
计算Coherence Score
常用的Coherence指标有两类,可根据需求选择:
- U_mass:取值范围为负无穷到0,数值越接近0主题一致性越高,仅依赖语料即可计算,速度快
- C_v:取值范围为0到1,数值越高主题一致性越高,结合词共现逻辑计算,结果更贴合人类感知
计算U_mass Coherence
# LSA的U_mass得分 lsa_umass = CoherenceModel( topics=lsa_topics, corpus=corpus, dictionary=dictionary, coherence='u_mass' ).get_coherence() # LDA的U_mass得分 lda_umass = CoherenceModel( topics=lda_topics, corpus=corpus, dictionary=dictionary, coherence='u_mass' ).get_coherence() print(f"LSA U_mass得分:{lsa_umass:.4f}") print(f"LDA U_mass得分:{lda_umass:.4f}")
计算C_v Coherence
# LSA的C_v得分 lsa_cv = CoherenceModel( topics=lsa_topics, texts=texts, dictionary=dictionary, coherence='c_v' ).get_coherence() # LDA的C_v得分 lda_cv = CoherenceModel( topics=lda_topics, texts=texts, dictionary=dictionary, coherence='c_v' ).get_coherence() print(f"LSA C_v得分:{lsa_cv:.4f}") print(f"LDA C_v得分:{lda_cv:.4f}")
注意:两类模型的得分对比需保证TOP_N_KEYWORDS取值、指标类型完全一致,结果才有可比性。
内容的提问来源于stack exchange,提问作者Aayush Sethi
相关产品推荐
相关产品推荐

