You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LDA模型对德语政治演讲数据集所有样本预测同一主题问题排查

问题描述

使用德语政治演讲数据集训练LDA模型,目标是将每篇演讲归类到对应主题中,但目前生成的主题相似度极高,且所有演讲都被预测为同一主题。尝试调整模型参数后无改善,以下是各环节代码:

预处理代码

def basic_preprocess_text(text):
    # Lowercase the text
    text = text.lower()

    # Replace umlauts and ß
    text = text.replace('ä', 'ae').replace('ö', 'oe').replace('ü', 'ue').replace('ß', 'ss')

    return text

stopwords = set(basic_preprocess_text(word) for word in stopwords)

def regex_filter(text) :
      # Remove everything before ":" if the pattern matches
    text = re.sub(r'.*\[.*\]:\s*', '', text)

    # Replacing poilitical term F.D.P as FDP
    text = re.sub(r'\b([A-Z]{1})\.?([A-Z]{1})\.?([A-Z]{1})\b', r'\1\2\3', text)

    # replace occurrences of "Nordrhein Westfalen" with "NRW"
    # text = re.sub(r'nordrhein[- .]?westfalen', 'NRW', text, flags=re.IGNORECASE)
    text = re.sub(r'nordrhein[- .]?westfalen', '', text, flags=re.IGNORECASE)

    #removing special charater
    text = re.sub(r'[^a-zA-Zäöüß]', ' ', text)

    return text

# Preprocessing function
def preprocess_text(text):

    text = regex_filter(text)

    # Tokenize the text
    tokens = word_tokenize(text, language='german')

    # Lemmatize the text & Basic preprocessing
    doc = nlp(' '.join(tokens))
    lemmatized_token = [basic_preprocess_text(token.lemma_) for token in doc if len(token.lemma_) > 2]

    # Remove stopwords
    lemmatized_text = ' '.join([word for word in lemmatized_token if word not in stopwords])

    return lemmatized_text
    #return tokens

Ngram处理代码

# Ensure 'Preprocessed_Speech' is treated as a list of tokens
df['Preprocessed_tokens'] = df['Preprocessed_Speech'].apply(lambda x: x.split())


# Create bigrams and trigrams
bigram = Phrases(df['Preprocessed_tokens'], min_count=3, threshold=5)
trigram = Phrases(bigram[df['Preprocessed_tokens']], min_count=3, threshold=5)

bigram_mod = Phraser(bigram)
trigram_mod = Phraser(trigram)

# Form Bigrams and Trigrams
texts = [trigram_mod[bigram_mod[text]] for text in df['Preprocessed_tokens']]

TF-IDF处理(曾尝试手动加权,目前未启用)

# Create TF-IDF matrix
texts_joined = [' '.join(text) for text in texts]
tfidf_vectorizer = TfidfVectorizer(max_df=0.50, min_df=5, ngram_range=(1, 2))
tfidf = tfidf_vectorizer.fit_transform(texts_joined)

# Convert TF-IDF matrix to a Gensim corpus
corpus = Sparse2Corpus(tfidf, documents_columns=False)

# Create the dictionary
id2word = {v: k for k, v in tfidf_vectorizer.vocabulary_.items()}

LDA训练代码

# Apply LDA with adjusted parameters
num_topics = 50  # Adjust number of topics for faster experimentation
passes = 20  # Increase number of passes
iterations = 1000  # Adjust number of iterations per pass
alpha = 'auto'  # Let gensim determine the optimal alpha
eta = 'auto'  # Let gensim determine the optimal eta

lda_model = models.LdaModel(corpus, num_topics=num_topics, id2word=id2word, passes=passes, iterations=iterations, alpha=alpha, eta=eta)

请问是操作存在错误,还是该数据集不适合LDA模型?


问题分析与解决方案

一、预处理环节的潜在问题

  1. 停用词过滤不彻底
    德语政治文本存在大量领域专属高频泛用词(如regierung、parlament),这类词未被过滤会直接主导主题分布,导致所有文档被归为同一主题。建议统计预处理后文本的词频,将前20-30个高频词加入停用词表;同时检查basic_preprocess_text对停用词的变音转换是否完全,避免因格式不一致漏过滤。

  2. 正则过滤的副作用
    正则r'.*\[.*\]:\s*'会匹配并删除所有包含方括号的前缀内容,可能误删大量有效演讲文本,导致文档只剩通用词汇。建议调整为精准匹配行首前缀:r'^\[.*\]:\s*';另外,删除Nordrhein Westfalen而非替换为NRW,会丢失地域主题信号,若数据集包含不同地区的演讲,会直接削弱主题区分度。

  3. 词形还原一致性问题
    需确保spaCy德语模型的词形还原输出,与自定义变音转换后的格式一致。比如spaCy输出的äußerung转换为aeusserung后,要保证停用词表中对应的也是该格式,避免漏过滤。

二、TF-IDF与LDA的兼容性问题

Gensim的LDA原生基于词频(BoW) corpus设计,TF-IDF转换会弱化主题的共现频率特征,可能导致模型无法有效区分主题。建议改用BoW corpus训练:

# 用Gensim原生工具构建词典和语料库
id2word = corpora.Dictionary(texts)
id2word.filter_extremes(no_below=5, no_above=0.5)
corpus = [id2word.doc2bow(text) for text in texts]

若坚持使用TF-IDF,需验证Sparse2Corpus的转换结果,确保维度匹配(scikit-learn的TF-IDF是文档×词汇,Gensim需要词汇×文档,documents_columns=False参数正确,但建议打印前几个样本确认)。

三、模型参数的不合理设置

  1. 主题数量过多
    设置num_topics=50远超政治演讲的天然主题范围(通常集中在内政、经济、外交等5-10大类),过多主题会导致模型无法收敛到清晰边界,反而生成高度相似的主题。先尝试将主题数降至5-10,观察区分度变化。

  2. alpha参数的自动推断陷阱
    alpha='auto'会让模型自动推断文档-主题的稀疏性,若数据集中文档主题分布本就集中,自动推断的alpha会强化这种趋势。尝试手动设置较小的固定值(如alpha=0.1),强制模型让文档分配到多个主题,避免单一主题主导。

  3. 迭代与遍历的有效性
    虽然设置了passes=20和iterations=1000,但模型可能在早期就已收敛,后续迭代无效。建议添加eval_every=1参数监控困惑度变化,当困惑度不再下降时停止训练。

四、数据集特性的影响

德语政治演讲本身可能存在以下特性,导致LDA效果不佳:

  • 主题重叠度高:政治演讲常涉及交叉主题(如经济政策关联就业、税收),若数据集中文档的主题分布无明显区分,LDA很难生成差异化主题。
  • 文本同质化:若数据集来自同一政党、同一时期的演讲,词汇和主题会高度相似,此时LDA无法区分主题。建议先分析数据集元数据(政党、演讲时间、议题标签),验证是否存在天然的主题区分度。

内容的提问来源于stack exchange,提问作者Ryu Ahmed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 11:34:54