新闻文本分类NMF模型特征提取异常:商业与科技类特征混淆
新闻分类NMF主题特征错位问题修复
问题概述
我正在开发新闻文档分类系统,将新闻分为**政治(politics)、商业(business)、科技(tech)、体育(sport)、娱乐(entertainment)**5个类别。通过NMF提取主题特征时出现异常:
- 本应属于科技类的特征被分配给主题0:
people, users, technology, said, use, internet, online, phone, mobile, using, digital, computer, net, information, many - 本应属于商业类的特征被分配给主题4:
said, company, year, growth, market, us, economy, shares, firm, 2004, bank, business, oil, rise, however
其余类别(体育、政治、娱乐)的主题特征匹配正确,需将商业与科技类的主题特征互换。
解决方案
方法1:手动修正主题-类别映射
直接在后续分类逻辑中手动调整主题与类别的对应关系,无需重新训练模型:
# 修正后的主题-类别映射 topic_to_category = { 0: 'tech', 1: 'sport', 2: 'politics', 3: 'entertainment', 4: 'business' } # 将预测主题转换为真实类别 def get_category(topic_idx): return topic_to_category[topic_idx]
方法2:优化模型参数,让主题自动匹配类别
通过调整预处理和NMF参数,提升主题特征的类别辨识度:
- 自定义停用词:过滤
said这类无意义高频词 - 调整TF-IDF参数:优化词汇过滤规则,保留更多行业专属特征
- 更换NMF求解器和损失函数:尝试更贴合分类场景的参数组合
修改后的代码示例:
# 自定义停用词集合 from nltk.corpus import stopwords custom_stopwords = set(stopwords.words('english')) custom_stopwords.update(['said', 'would', 'people']) # 更新文本清理函数 def clean_white_spaces(text): text = text.lower().replace('\n', ' ').replace('\r', '').strip() text = re.sub(' +', ' ', text) text = re.sub(r'[^\w\s]', '', text) tokens = word_tokenize(text) tokens = [w for w in tokens if not w in custom_stopwords] return ' '.join(tokens) # 重新训练TF-IDF vectorizer = TfidfVectorizer(encoding='utf-8', ngram_range=(1,3), stop_words=custom_stopwords, lowercase=False, max_df=0.9, min_df=5, norm='l2', sublinear_tf=True) X = vectorizer.fit_transform(X_train).toarray() # 重新训练NMF nmf = NMF(n_components=5, random_state=42, solver="cd", beta_loss="frobenius", max_iter=1000) model = nmf.fit(X) # 查看优化后的主题特征 display_topics(nmf, vectorizer.get_feature_names_out(), 15)
内容的提问来源于stack exchange,提问作者Shivam
相关产品推荐
相关产品推荐

