You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Gensim 3.8.3的LdaMallet模型遇AttributeError问题求助

LdaMallet文档主题提取函数的报错修正方案

运行代码时出现AttributeError: 'LdaMallet' object has no attribute 'per_word_topics',原因是原代码适配的是Gensim内置的LdaModel(该模型拥有per_word_topics属性),而你使用的LdaMallet模型没有这个属性,且返回的文档主题数据格式也和LdaModel不同。

以下是修正后的完整代码:

import pandas as pd

def format_topics_sentences(ldamodel=lda_model, corpus=corpus, texts=data1):
    # 初始化输出列表
    final = []
    # 遍历每个文档的主题结果
    for i, row in enumerate(ldamodel[corpus]):
        # LdaMallet直接返回主题(编号, 概率)的元组列表,直接排序即可
        row = sorted(row, key=lambda x: (x[1]), reverse=True)
        # 获取每个文档的主导主题、贡献占比和关键词
        for j, (topic_num, prop_topic) in enumerate(row):
            if j == 0:  # 取概率最高的主导主题
                wp = ldamodel.show_topic(topic_num)
                topic_keywords = ", ".join([word for word, prop in wp])
                lists1 = int(topic_num), round(prop_topic, 4), topic_keywords
                final.append(lists1)
            else:
                break
    # 转换为DataFrame并合并原文内容
    sent_topics_df = pd.DataFrame(final, columns=['Dominant_Topic', 'Perc_Contribution', 'Topic_Keywords'])
    contents = pd.Series(texts)
    sent_topics_df = pd.concat([sent_topics_df, contents], axis=1)
    
    return sent_topics_df

# 调用函数(参数保持你的配置)
df_topic_sents_keywords = format_topics_sentences(ldamodel=optimal_model, corpus=corpus, texts=texts)

# 格式化输出结果
df_dominant_topic = df_topic_sents_keywords.reset_index()
df_dominant_topic.columns = ['Document_No', 'Dominant_Topic', 'Topic_Perc_Contrib', 'Keywords', 'Text']

# 查看前10条数据
df_dominant_topic.head(10)

关键修改说明

  • 移除了原代码中针对per_word_topics的判断逻辑,因为LdaMallet调用ldamodel[corpus]时,直接返回的就是每个文档的主题(编号, 概率)元组列表,无需额外提取row_list[0]。
  • 保持了原有的主题排序、主导主题提取和结果合并逻辑,确保功能和预期一致。

内容的提问来源于stack exchange,提问作者nazeli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 02:05:36