You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的Scikit-learn绘制TF-IDF SVM文本分类的特征与超平面

如何绘制TF-IDF特征与SVM分类超平面?

你好呀!我明白你已经搞定了TF-IDF+SVM的文本分类,现在想把特征和分类超平面可视化出来——这确实是个能帮你直观理解模型的好想法。不过有个关键点得先说明:原始的TF-IDF特征是高维的(通常几千甚至上万维),没法直接在平面上绘制,所以我们得先把特征降维到2维,再结合SVM的决策边界来展示。

下面一步步来实现:

核心思路

  1. 用降维算法(比如PCA或t-SNE)把高维TF-IDF特征压缩到2维,这样就能在平面上展示样本分布
  2. 在降维后的2维空间里训练一个SVM模型,绘制它的决策边界(也就是你说的超平面)
  3. 把降维后的样本点和决策边界画在同一张图里

修改后的完整代码

我基于你现有的代码做了调整,加入了降维和绘图逻辑:

import os
import numpy as np
import matplotlib.pyplot as plt
from sklearn.svm import LinearSVC
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.decomposition import PCA
# 如果想用非线性降维,替换成下面的TSNE
# from sklearn.manifold import TSNE

def make_Corpus(root_dir):
    polarity_dirs = [os.path.join(root_dir,f) for f in os.listdir(root_dir)]
    corpus = []
    for polarity_dir in polarity_dirs:
        reviews = [os.path.join(polarity_dir,f) for f in os.listdir(polarity_dir)]
        for review in reviews:
            doc_string = ""
            with open(review) as rev:
                for line in rev:
                    doc_string += line
            corpus.append(doc_string)
    return corpus

# 加载数据集
root_dir = 'txt_sentoken'
corpus = make_Corpus(root_dir)
labels = np.zeros(2000)
labels[0:1000] = 0  # 负样本
labels[1000:2000] = 1  # 正样本

# 生成TF-IDF特征
vectorizer = TfidfVectorizer(min_df=5, max_df=0.8, sublinear_tf=True, use_idf=True, stop_words='english')
tf_idf_matrix = vectorizer.fit_transform(corpus)

# 降维到2维:这里用PCA(线性降维,速度快),也可以换成TSNE(非线性,对文本分布更友好但速度慢)
# -------------------
# 选项1:PCA降维
pca = PCA(n_components=2, random_state=42)
tf_idf_2d = pca.fit_transform(tf_idf_matrix.toarray())

# 选项2:TSNE降维(取消注释即可)
# tsne = TSNE(n_components=2, random_state=42, perplexity=30)
# tf_idf_2d = tsne.fit_transform(tf_idf_matrix.toarray())
# -------------------

# 在降维后的2维空间训练SVM,用于绘制决策边界
svm_2d_model = LinearSVC()
svm_2d_model.fit(tf_idf_2d, labels)

# 开始绘图
plt.figure(figsize=(12, 8))

# 1. 绘制降维后的样本散点图
plt.scatter(tf_idf_2d[labels == 0, 0], tf_idf_2d[labels == 0, 1], 
            label='Negative Reviews', alpha=0.6, color='#1f77b4')
plt.scatter(tf_idf_2d[labels == 1, 0], tf_idf_2d[labels == 1, 1], 
            label='Positive Reviews', alpha=0.6, color='#ff7f0e')

# 2. 绘制SVM的决策边界(超平面)
# 创建网格点覆盖整个绘图区域
x_min, x_max = tf_idf_2d[:, 0].min() - 0.5, tf_idf_2d[:, 0].max() + 0.5
y_min, y_max = tf_idf_2d[:, 1].min() - 0.5, tf_idf_2d[:, 1].max() + 0.5
xx, yy = np.meshgrid(np.arange(x_min, x_max, 0.02),
                     np.arange(y_min, y_max, 0.02))

# 预测网格点的类别,生成决策边界
Z = svm_2d_model.predict(np.c_[xx.ravel(), yy.ravel()])
Z = Z.reshape(xx.shape)
plt.contourf(xx, yy, Z, alpha=0.2, cmap=plt.cm.coolwarm)

# 设置图的标签和标题
plt.xlabel('Dimension 1')
plt.ylabel('Dimension 2')
plt.title('TF-IDF Features (2D Reduced) & SVM Classification Boundary')
plt.legend()
plt.grid(True, alpha=0.3)
plt.show()

关键细节说明

  1. 降维选择:

    • PCA是线性降维,计算速度快,适合快速看整体分布,但可能会丢失一些非线性的特征关联
    • t-SNE是非线性降维,能更好地保留样本间的局部结构,对文本这种复杂数据的可视化效果通常更好,但计算时间更长(处理2000个样本大概需要几十秒)
  2. 超平面绘制的小技巧:
    原始的SVM是在高维TF-IDF空间训练的,没法直接映射到2维空间画超平面。所以我们用降维后的2维特征重新训练了一个SVM,这样就能直接在2维网格上生成决策边界——虽然这和原始高维模型不完全一样,但足够帮你直观理解分类逻辑。

  3. 单独查看单个词的TF-IDF分布:
    如果你想看看某个特定词(比如"good")的TF-IDF值分布,可以这样做:

    # 获取目标词在词汇表中的索引
    target_word = "good"
    word_idx = vectorizer.vocabulary_.get(target_word)
    if word_idx is not None:
        # 提取该词的TF-IDF值
        word_tfidf = tf_idf_matrix[:, word_idx].toarray().flatten()
        # 绘制箱线图,对比正负样本的TF-IDF值
        plt.boxplot([word_tfidf[labels==0], word_tfidf[labels==1]], labels=['Negative', 'Positive'])
        plt.title(f'TF-IDF Value of "{target_word}" in Reviews')
        plt.show()
    else:
        print(f"Word '{target_word}' not found in vocabulary!")
    

内容的提问来源于stack exchange,提问作者OnePunchMan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:51:26