You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

新闻文本分类NMF模型特征提取异常:商业与科技类特征混淆

新闻分类NMF主题特征错位问题修复

问题概述

我正在开发新闻文档分类系统,将新闻分为**政治(politics)、商业(business)、科技(tech)、体育(sport)、娱乐(entertainment)**5个类别。通过NMF提取主题特征时出现异常:

  • 本应属于科技类的特征被分配给主题0:people, users, technology, said, use, internet, online, phone, mobile, using, digital, computer, net, information, many
  • 本应属于商业类的特征被分配给主题4:said, company, year, growth, market, us, economy, shares, firm, 2004, bank, business, oil, rise, however
    其余类别(体育、政治、娱乐)的主题特征匹配正确,需将商业与科技类的主题特征互换。

解决方案

方法1:手动修正主题-类别映射

直接在后续分类逻辑中手动调整主题与类别的对应关系,无需重新训练模型:

# 修正后的主题-类别映射
topic_to_category = {
    0: 'tech',
    1: 'sport',
    2: 'politics',
    3: 'entertainment',
    4: 'business'
}

# 将预测主题转换为真实类别
def get_category(topic_idx):
    return topic_to_category[topic_idx]

方法2:优化模型参数,让主题自动匹配类别

通过调整预处理和NMF参数,提升主题特征的类别辨识度:

  • 自定义停用词:过滤said这类无意义高频词
  • 调整TF-IDF参数:优化词汇过滤规则,保留更多行业专属特征
  • 更换NMF求解器和损失函数:尝试更贴合分类场景的参数组合

修改后的代码示例:

# 自定义停用词集合
from nltk.corpus import stopwords
custom_stopwords = set(stopwords.words('english'))
custom_stopwords.update(['said', 'would', 'people'])

# 更新文本清理函数
def clean_white_spaces(text):
    text = text.lower().replace('\n', ' ').replace('\r', '').strip()
    text = re.sub(' +', ' ', text)
    text = re.sub(r'[^\w\s]', '', text)
    tokens = word_tokenize(text)
    tokens = [w for w in tokens if not w in custom_stopwords]
    return ' '.join(tokens)

# 重新训练TF-IDF
vectorizer = TfidfVectorizer(encoding='utf-8',
                             ngram_range=(1,3),
                             stop_words=custom_stopwords,
                             lowercase=False,
                             max_df=0.9,
                             min_df=5,
                             norm='l2',
                             sublinear_tf=True)
X = vectorizer.fit_transform(X_train).toarray()

# 重新训练NMF
nmf = NMF(n_components=5, random_state=42, solver="cd", beta_loss="frobenius", max_iter=1000)
model = nmf.fit(X)

# 查看优化后的主题特征
display_topics(nmf, vectorizer.get_feature_names_out(), 15)

内容的提问来源于stack exchange,提问作者Shivam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 18:24:57