Gensim LDA提取文档主题分数时元组排序报错求助
问题解决:TypeError: '<' not supported between instances of 'tuple' and 'int'
错误原因
报错核心是ldamodel[corpus]返回的row结构不符合预期:正常情况下每个row应为文档的主题分布列表(元素是(topic_id, 概率值)的元组),但实际返回的row是嵌套结构(比如外层为元组,主题分布仅存于row[0]中),导致排序时尝试比较tuple和int类型的值,触发类型错误。
修复后的代码
import pandas as pd def format_topics_sentences(ldamodel=optimal_model, corpus=corpus, texts=data): # 初始化输出DataFrame sent_topics_df = pd.DataFrame() # 获取每个文档的主导主题 for i, row in enumerate(ldamodel[corpus]): # 处理嵌套结构:如果row是元组,提取第一个元素作为主题分布列表 if isinstance(row, tuple): row = row[0] # 按主题概率降序排序 row_sorted = sorted(row, key=lambda x: x[1], reverse=True) # 提取主导主题、贡献占比和关键词 if row_sorted: # 避免空主题分布导致索引错误 topic_num, prop_topic = row_sorted[0] wp = ldamodel.show_topic(topic_num) topic_keywords = ", ".join([word for word, prop in wp]) # 用pd.concat替代弃用的append new_row = pd.Series([int(topic_num), round(prop_topic,4), topic_keywords], index=['Dominant_Topic', 'Perc_Contribution', 'Topic_Keywords']) sent_topics_df = pd.concat([sent_topics_df, new_row.to_frame().T], ignore_index=True) # 添加原始文本 contents = pd.Series(texts, name='Text') sent_topics_df = pd.concat([sent_topics_df, contents], axis=1) return sent_topics_df # 生成结果并格式化 df_topic_sents_keywords = format_topics_sentences(ldamodel=optimal_model, corpus=corpus, texts=data) df_dominant_topic = df_topic_sents_keywords.reset_index() df_dominant_topic.columns = ['Document_No', 'Dominant_Topic', 'Topic_Perc_Contrib', 'Keywords', 'Text'] # 展示前10条 print(df_dominant_topic.head(10))
关键修复点
- 增加结构判断:自动识别并处理嵌套的元组结构,确保排序对象是正确的主题分布列表。
- 简化排序逻辑:直接用主题概率值作为排序键,避免不必要的元组嵌套。
- 替换弃用方法:完全用
pd.concat替代df.append,适配Pandas新版本规范。 - 增加非空校验:防止空主题分布导致的索引越界错误。
内容的提问来源于stack exchange,提问作者埃塞ABELA
相关产品推荐
相关产品推荐

