LDA模型对德语政治演讲数据集所有样本预测同一主题问题排查
使用德语政治演讲数据集训练LDA模型,目标是将每篇演讲归类到对应主题中,但目前生成的主题相似度极高,且所有演讲都被预测为同一主题。尝试调整模型参数后无改善,以下是各环节代码:
预处理代码
def basic_preprocess_text(text): # Lowercase the text text = text.lower() # Replace umlauts and ß text = text.replace('ä', 'ae').replace('ö', 'oe').replace('ü', 'ue').replace('ß', 'ss') return text stopwords = set(basic_preprocess_text(word) for word in stopwords) def regex_filter(text) : # Remove everything before ":" if the pattern matches text = re.sub(r'.*\[.*\]:\s*', '', text) # Replacing poilitical term F.D.P as FDP text = re.sub(r'\b([A-Z]{1})\.?([A-Z]{1})\.?([A-Z]{1})\b', r'\1\2\3', text) # replace occurrences of "Nordrhein Westfalen" with "NRW" # text = re.sub(r'nordrhein[- .]?westfalen', 'NRW', text, flags=re.IGNORECASE) text = re.sub(r'nordrhein[- .]?westfalen', '', text, flags=re.IGNORECASE) #removing special charater text = re.sub(r'[^a-zA-Zäöüß]', ' ', text) return text # Preprocessing function def preprocess_text(text): text = regex_filter(text) # Tokenize the text tokens = word_tokenize(text, language='german') # Lemmatize the text & Basic preprocessing doc = nlp(' '.join(tokens)) lemmatized_token = [basic_preprocess_text(token.lemma_) for token in doc if len(token.lemma_) > 2] # Remove stopwords lemmatized_text = ' '.join([word for word in lemmatized_token if word not in stopwords]) return lemmatized_text #return tokens
Ngram处理代码
# Ensure 'Preprocessed_Speech' is treated as a list of tokens df['Preprocessed_tokens'] = df['Preprocessed_Speech'].apply(lambda x: x.split()) # Create bigrams and trigrams bigram = Phrases(df['Preprocessed_tokens'], min_count=3, threshold=5) trigram = Phrases(bigram[df['Preprocessed_tokens']], min_count=3, threshold=5) bigram_mod = Phraser(bigram) trigram_mod = Phraser(trigram) # Form Bigrams and Trigrams texts = [trigram_mod[bigram_mod[text]] for text in df['Preprocessed_tokens']]
TF-IDF处理(曾尝试手动加权,目前未启用)
# Create TF-IDF matrix texts_joined = [' '.join(text) for text in texts] tfidf_vectorizer = TfidfVectorizer(max_df=0.50, min_df=5, ngram_range=(1, 2)) tfidf = tfidf_vectorizer.fit_transform(texts_joined) # Convert TF-IDF matrix to a Gensim corpus corpus = Sparse2Corpus(tfidf, documents_columns=False) # Create the dictionary id2word = {v: k for k, v in tfidf_vectorizer.vocabulary_.items()}
LDA训练代码
# Apply LDA with adjusted parameters num_topics = 50 # Adjust number of topics for faster experimentation passes = 20 # Increase number of passes iterations = 1000 # Adjust number of iterations per pass alpha = 'auto' # Let gensim determine the optimal alpha eta = 'auto' # Let gensim determine the optimal eta lda_model = models.LdaModel(corpus, num_topics=num_topics, id2word=id2word, passes=passes, iterations=iterations, alpha=alpha, eta=eta)
请问是操作存在错误,还是该数据集不适合LDA模型?
一、预处理环节的潜在问题
停用词过滤不彻底
德语政治文本存在大量领域专属高频泛用词(如regierung、parlament),这类词未被过滤会直接主导主题分布,导致所有文档被归为同一主题。建议统计预处理后文本的词频,将前20-30个高频词加入停用词表;同时检查basic_preprocess_text对停用词的变音转换是否完全,避免因格式不一致漏过滤。正则过滤的副作用
正则r'.*\[.*\]:\s*'会匹配并删除所有包含方括号的前缀内容,可能误删大量有效演讲文本,导致文档只剩通用词汇。建议调整为精准匹配行首前缀:r'^\[.*\]:\s*';另外,删除Nordrhein Westfalen而非替换为NRW,会丢失地域主题信号,若数据集包含不同地区的演讲,会直接削弱主题区分度。词形还原一致性问题
需确保spaCy德语模型的词形还原输出,与自定义变音转换后的格式一致。比如spaCy输出的äußerung转换为aeusserung后,要保证停用词表中对应的也是该格式,避免漏过滤。
二、TF-IDF与LDA的兼容性问题
Gensim的LDA原生基于词频(BoW) corpus设计,TF-IDF转换会弱化主题的共现频率特征,可能导致模型无法有效区分主题。建议改用BoW corpus训练:
# 用Gensim原生工具构建词典和语料库 id2word = corpora.Dictionary(texts) id2word.filter_extremes(no_below=5, no_above=0.5) corpus = [id2word.doc2bow(text) for text in texts]
若坚持使用TF-IDF,需验证Sparse2Corpus的转换结果,确保维度匹配(scikit-learn的TF-IDF是文档×词汇,Gensim需要词汇×文档,documents_columns=False参数正确,但建议打印前几个样本确认)。
三、模型参数的不合理设置
主题数量过多
设置num_topics=50远超政治演讲的天然主题范围(通常集中在内政、经济、外交等5-10大类),过多主题会导致模型无法收敛到清晰边界,反而生成高度相似的主题。先尝试将主题数降至5-10,观察区分度变化。alpha参数的自动推断陷阱
alpha='auto'会让模型自动推断文档-主题的稀疏性,若数据集中文档主题分布本就集中,自动推断的alpha会强化这种趋势。尝试手动设置较小的固定值(如alpha=0.1),强制模型让文档分配到多个主题,避免单一主题主导。迭代与遍历的有效性
虽然设置了passes=20和iterations=1000,但模型可能在早期就已收敛,后续迭代无效。建议添加eval_every=1参数监控困惑度变化,当困惑度不再下降时停止训练。
四、数据集特性的影响
德语政治演讲本身可能存在以下特性,导致LDA效果不佳:
- 主题重叠度高:政治演讲常涉及交叉主题(如经济政策关联就业、税收),若数据集中文档的主题分布无明显区分,LDA很难生成差异化主题。
- 文本同质化:若数据集来自同一政党、同一时期的演讲,词汇和主题会高度相似,此时LDA无法区分主题。建议先分析数据集元数据(政党、演讲时间、议题标签),验证是否存在天然的主题区分度。
内容的提问来源于stack exchange,提问作者Ryu Ahmed

