You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

运行Gensim的CoherenceModel计算c_v分数出现KeyError如何解决

问题根因

你遇到的报错核心原因是:topics列表中存在未被收录进你基于data_words构建的词典id2word的词汇(如报错信息中的afgelopen、lamp),CoherenceModel在计算c_v分数时需要将主题词映射为词典ID,找不到对应映射就会触发KeyError。

解决方案

你可以选择以下任意一种方案解决问题:

  • 方案1:过滤主题列表中的非法词汇
    在完成词典构建后,先对topics做过滤,仅保留存在于id2word中的词汇再传入CoherenceModel:
import gensim.corpora as corpora
from gensim.models.coherencemodel import CoherenceModel

id2word = corpora.Dictionary(data_words)
corpus = [id2word.doc2bow(text) for text in data_words]

# 新增过滤逻辑
valid_tokens = set(id2word.token2id.keys())
filtered_topics = []
for topic in topics:
    # 仅保留词典中存在的词
    valid_topic = [token for token in topic if token in valid_tokens]
    # 仅保留非空的主题,可根据业务需求调整空主题的处理逻辑
    if len(valid_topic) > 0:
        filtered_topics.append(valid_topic)

coherence_score = CoherenceModel(topics=filtered_topics,
                          texts = data_words, 
                          corpus= corpus, 
                          dictionary= id2word, 
                          coherence= 'c_v', 
                          topn=20).get_coherence()
  • 方案2:扩展词典构建的语料范围
    如果需要保留所有主题词,可以在构建词典时将topics的内容也纳入语料,保证所有主题词都被收录:
import gensim.corpora as corpora
from gensim.models.coherencemodel import CoherenceModel

# 合并语料与主题词列表构建词典
all_texts = data_words + topics
id2word = corpora.Dictionary(all_texts)
corpus = [id2word.doc2bow(text) for text in data_words]

coherence_score = CoherenceModel(topics=topics,
                          texts = data_words, 
                          corpus= corpus, 
                          dictionary= id2word, 
                          coherence= 'c_v', 
                          topn=20).get_coherence()

注意:如果使用方案2,需要保证data_words中存在足够多的主题词共现样本,否则计算出的c_v分数参考价值会下降。

内容的提问来源于stack exchange,提问作者Emil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 01:24:04