使用Gensim 3.8.3的LdaMallet模型遇AttributeError问题求助
LdaMallet文档主题提取函数的报错修正方案
运行代码时出现AttributeError: 'LdaMallet' object has no attribute 'per_word_topics',原因是原代码适配的是Gensim内置的LdaModel(该模型拥有per_word_topics属性),而你使用的LdaMallet模型没有这个属性,且返回的文档主题数据格式也和LdaModel不同。
以下是修正后的完整代码:
import pandas as pd def format_topics_sentences(ldamodel=lda_model, corpus=corpus, texts=data1): # 初始化输出列表 final = [] # 遍历每个文档的主题结果 for i, row in enumerate(ldamodel[corpus]): # LdaMallet直接返回主题(编号, 概率)的元组列表,直接排序即可 row = sorted(row, key=lambda x: (x[1]), reverse=True) # 获取每个文档的主导主题、贡献占比和关键词 for j, (topic_num, prop_topic) in enumerate(row): if j == 0: # 取概率最高的主导主题 wp = ldamodel.show_topic(topic_num) topic_keywords = ", ".join([word for word, prop in wp]) lists1 = int(topic_num), round(prop_topic, 4), topic_keywords final.append(lists1) else: break # 转换为DataFrame并合并原文内容 sent_topics_df = pd.DataFrame(final, columns=['Dominant_Topic', 'Perc_Contribution', 'Topic_Keywords']) contents = pd.Series(texts) sent_topics_df = pd.concat([sent_topics_df, contents], axis=1) return sent_topics_df # 调用函数(参数保持你的配置) df_topic_sents_keywords = format_topics_sentences(ldamodel=optimal_model, corpus=corpus, texts=texts) # 格式化输出结果 df_dominant_topic = df_topic_sents_keywords.reset_index() df_dominant_topic.columns = ['Document_No', 'Dominant_Topic', 'Topic_Perc_Contrib', 'Keywords', 'Text'] # 查看前10条数据 df_dominant_topic.head(10)
关键修改说明
- 移除了原代码中针对
per_word_topics的判断逻辑,因为LdaMallet调用ldamodel[corpus]时,直接返回的就是每个文档的主题(编号, 概率)元组列表,无需额外提取row_list[0]。 - 保持了原有的主题排序、主导主题提取和结果合并逻辑,确保功能和预期一致。
内容的提问来源于stack exchange,提问作者nazeli
相关产品推荐
相关产品推荐

