运行Gensim的CoherenceModel计算c_v分数出现KeyError如何解决
问题根因
你遇到的报错核心原因是:topics列表中存在未被收录进你基于data_words构建的词典id2word的词汇(如报错信息中的afgelopen、lamp),CoherenceModel在计算c_v分数时需要将主题词映射为词典ID,找不到对应映射就会触发KeyError。
解决方案
你可以选择以下任意一种方案解决问题:
- 方案1:过滤主题列表中的非法词汇
在完成词典构建后,先对topics做过滤,仅保留存在于id2word中的词汇再传入CoherenceModel:
import gensim.corpora as corpora from gensim.models.coherencemodel import CoherenceModel id2word = corpora.Dictionary(data_words) corpus = [id2word.doc2bow(text) for text in data_words] # 新增过滤逻辑 valid_tokens = set(id2word.token2id.keys()) filtered_topics = [] for topic in topics: # 仅保留词典中存在的词 valid_topic = [token for token in topic if token in valid_tokens] # 仅保留非空的主题,可根据业务需求调整空主题的处理逻辑 if len(valid_topic) > 0: filtered_topics.append(valid_topic) coherence_score = CoherenceModel(topics=filtered_topics, texts = data_words, corpus= corpus, dictionary= id2word, coherence= 'c_v', topn=20).get_coherence()
- 方案2:扩展词典构建的语料范围
如果需要保留所有主题词,可以在构建词典时将topics的内容也纳入语料,保证所有主题词都被收录:
import gensim.corpora as corpora from gensim.models.coherencemodel import CoherenceModel # 合并语料与主题词列表构建词典 all_texts = data_words + topics id2word = corpora.Dictionary(all_texts) corpus = [id2word.doc2bow(text) for text in data_words] coherence_score = CoherenceModel(topics=topics, texts = data_words, corpus= corpus, dictionary= id2word, coherence= 'c_v', topn=20).get_coherence()
注意:如果使用方案2,需要保证data_words中存在足够多的主题词共现样本,否则计算出的c_v分数参考价值会下降。
内容的提问来源于stack exchange,提问作者Emil
相关产品推荐
相关产品推荐

