You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于判别特征的文本分类:Sklearn实现方法咨询

好的,我来帮你解决这个问题——用Sklearn实现基于判别特征(某类占比≥90%的词/n-gram)的文本分类,核心就是先按类别筛选符合要求的特征,再用这些特征训练模型。下面是具体的实现步骤和代码:

实现思路

核心逻辑拆解:

  • 先提取所有候选特征(单字、多字词或n-gram)
  • 计算每个特征在各个类别中的占比(即该特征的所有出现实例中,属于某类的比例)
  • 筛选出至少在一个类别中占比≥90%的特征
  • 用筛选后的特征训练分类器(避免无关特征干扰)
具体代码实现

1. 导入依赖库

import numpy as np
from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report

2. 准备数据(替换成你的数据集即可)

# 示例文本数据
texts = [
    "苹果发布最新款iPhone,搭载A17芯片",
    "华为Mate 60系列支持卫星通信,性能强劲",
    "iPhone 15的摄像头升级明显,拍照更清晰",
    "华为新机型采用麒麟芯片,续航能力出色",
    "苹果官网开启iPhone 15预购,销量火爆",
    "华为鸿蒙系统更新,新增多项功能",
    "微软发布Windows 11新版本,优化用户体验",
    "Surface Pro 9搭载最新处理器,办公效率高",
    "微软Edge浏览器新增AI功能,智能助手更贴心",
    "Surface Laptop 5轻薄便携,适合移动办公"
]
# 对应的类别标签:0=苹果, 1=华为, 2=微软
labels = [0,1,0,1,0,1,2,2,2,2]

3. 提取初始特征(支持词或n-gram)

这里用CountVectorizer,你也可以换成TfidfVectorizer,逻辑完全一致:

# 初始化向量器,ngram_range=(1,2)表示提取单字+双字词,可按需调整
vectorizer = CountVectorizer(ngram_range=(1,2), stop_words=None)
# 拟合训练文本,得到词频矩阵X(形状:[样本数, 特征数])
X = vectorizer.fit_transform(texts)
# 获取所有特征的名称(词/n-gram)
feature_names = vectorizer.get_feature_names_out()
# 拆分训练集和测试集(必须用训练集做特征筛选,避免数据泄露)
X_train, X_test, y_train, y_test = train_test_split(X, labels, test_size=0.2, random_state=42)

4. 按类别计算特征占比并筛选判别特征

这里我们计算每个特征在某类中的总词频占该特征总词频的比例,筛选出至少有一个类别占比≥90%的特征:

# 获取训练集中的所有类别
unique_classes = np.unique(y_train)
# 计算每个类别下,所有特征的总词频
class_tf = []
for cls in unique_classes:
    # 取出该类所有样本的词频矩阵,求和得到该类每个特征的总出现次数
    cls_tf = X_train[y_train == cls].sum(axis=0).A1
    class_tf.append(cls_tf)
class_tf = np.array(class_tf)

# 计算每个特征在所有类别中的总词频
total_tf = class_tf.sum(axis=0)
# 避免除以0的情况,把总词频为0的特征设为1(这类特征本来就不会被选中)
total_tf[total_tf == 0] = 1

# 计算每个特征在各个类别中的占比
class_ratio = class_tf / total_tf

# 筛选特征:至少在一个类别中占比≥90%
selected_mask = (class_ratio >= 0.9).any(axis=0)
# 获取选中的特征和对应的索引
selected_features = feature_names[selected_mask]
selected_indices = np.where(selected_mask)[0]

# 过滤训练集和测试集,只保留符合要求的判别特征
X_train_selected = X_train[:, selected_indices]
X_test_selected = X_test[:, selected_indices]

print(f"原始特征数:{len(feature_names)},筛选后特征数:{len(selected_features)}")
print("选中的判别特征示例:", selected_features[:10])

可选:按文档出现比例筛选特征

如果你的需求是特征出现在某类文档中的比例≥90%(即该特征出现的所有文档里,90%属于某类),可以修改这部分代码:

# 改用binary=True的向量器,统计特征是否在文档中出现(而非词频)
vectorizer_binary = CountVectorizer(ngram_range=(1,2), binary=True)
X_binary = vectorizer_binary.fit_transform(texts)
X_train_bin, X_test_bin, y_train_bin, y_test_bin = train_test_split(X_binary, labels, test_size=0.2, random_state=42)

# 计算每个特征在各类别中的文档出现次数
class_doc_count = []
for cls in unique_classes:
    cls_docs = X_train_bin[y_train_bin == cls].sum(axis=0).A1
    class_doc_count.append(cls_docs)
class_doc_count = np.array(class_doc_count)

# 计算每个特征出现的总文档数
total_doc_count = class_doc_count.sum(axis=0)
total_doc_count[total_doc_count == 0] = 1

# 计算每个特征在各类别中的文档占比
class_doc_ratio = class_doc_count / total_doc_count

# 筛选特征
selected_mask_bin = (class_doc_ratio >= 0.9).any(axis=0)
selected_features_bin = vectorizer_binary.get_feature_names_out()[selected_mask_bin]
selected_indices_bin = np.where(selected_mask_bin)[0]

X_train_selected_bin = X_train_bin[:, selected_indices_bin]
X_test_selected_bin = X_test_bin[:, selected_indices_bin]

5. 用筛选后的特征训练分类器

# 初始化分类器(这里用逻辑回归,你可以换成SVM、朴素贝叶斯等)
clf = LogisticRegression()
# 用筛选后的特征训练模型
clf.fit(X_train_selected, y_train)
# 预测测试集
y_pred = clf.predict(X_test_selected)
# 输出分类评估报告
print(classification_report(y_test, y_pred))
关键注意事项
  • 必须基于训练集计算特征占比,绝对不能用测试集的数据参与特征选择,否则会导致数据泄露,模型泛化能力下降。
  • 如果你想用TfidfVectorizer,只需要把初始特征提取部分替换成TfidfVectorizer,后续的占比计算逻辑完全一致。
  • 可以根据需求调整ngram_range(比如(1,3)提取三字词)、占比阈值(比如90%改成85%)等参数。

内容的提问来源于stack exchange,提问作者Dorin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:05:15