如何用Python的Scikit-learn绘制TF-IDF SVM文本分类的特征与超平面
如何绘制TF-IDF特征与SVM分类超平面?
你好呀!我明白你已经搞定了TF-IDF+SVM的文本分类,现在想把特征和分类超平面可视化出来——这确实是个能帮你直观理解模型的好想法。不过有个关键点得先说明:原始的TF-IDF特征是高维的(通常几千甚至上万维),没法直接在平面上绘制,所以我们得先把特征降维到2维,再结合SVM的决策边界来展示。
下面一步步来实现:
核心思路
- 用降维算法(比如PCA或t-SNE)把高维TF-IDF特征压缩到2维,这样就能在平面上展示样本分布
- 在降维后的2维空间里训练一个SVM模型,绘制它的决策边界(也就是你说的超平面)
- 把降维后的样本点和决策边界画在同一张图里
修改后的完整代码
我基于你现有的代码做了调整,加入了降维和绘图逻辑:
import os import numpy as np import matplotlib.pyplot as plt from sklearn.svm import LinearSVC from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.decomposition import PCA # 如果想用非线性降维,替换成下面的TSNE # from sklearn.manifold import TSNE def make_Corpus(root_dir): polarity_dirs = [os.path.join(root_dir,f) for f in os.listdir(root_dir)] corpus = [] for polarity_dir in polarity_dirs: reviews = [os.path.join(polarity_dir,f) for f in os.listdir(polarity_dir)] for review in reviews: doc_string = "" with open(review) as rev: for line in rev: doc_string += line corpus.append(doc_string) return corpus # 加载数据集 root_dir = 'txt_sentoken' corpus = make_Corpus(root_dir) labels = np.zeros(2000) labels[0:1000] = 0 # 负样本 labels[1000:2000] = 1 # 正样本 # 生成TF-IDF特征 vectorizer = TfidfVectorizer(min_df=5, max_df=0.8, sublinear_tf=True, use_idf=True, stop_words='english') tf_idf_matrix = vectorizer.fit_transform(corpus) # 降维到2维:这里用PCA(线性降维,速度快),也可以换成TSNE(非线性,对文本分布更友好但速度慢) # ------------------- # 选项1:PCA降维 pca = PCA(n_components=2, random_state=42) tf_idf_2d = pca.fit_transform(tf_idf_matrix.toarray()) # 选项2:TSNE降维(取消注释即可) # tsne = TSNE(n_components=2, random_state=42, perplexity=30) # tf_idf_2d = tsne.fit_transform(tf_idf_matrix.toarray()) # ------------------- # 在降维后的2维空间训练SVM,用于绘制决策边界 svm_2d_model = LinearSVC() svm_2d_model.fit(tf_idf_2d, labels) # 开始绘图 plt.figure(figsize=(12, 8)) # 1. 绘制降维后的样本散点图 plt.scatter(tf_idf_2d[labels == 0, 0], tf_idf_2d[labels == 0, 1], label='Negative Reviews', alpha=0.6, color='#1f77b4') plt.scatter(tf_idf_2d[labels == 1, 0], tf_idf_2d[labels == 1, 1], label='Positive Reviews', alpha=0.6, color='#ff7f0e') # 2. 绘制SVM的决策边界(超平面) # 创建网格点覆盖整个绘图区域 x_min, x_max = tf_idf_2d[:, 0].min() - 0.5, tf_idf_2d[:, 0].max() + 0.5 y_min, y_max = tf_idf_2d[:, 1].min() - 0.5, tf_idf_2d[:, 1].max() + 0.5 xx, yy = np.meshgrid(np.arange(x_min, x_max, 0.02), np.arange(y_min, y_max, 0.02)) # 预测网格点的类别,生成决策边界 Z = svm_2d_model.predict(np.c_[xx.ravel(), yy.ravel()]) Z = Z.reshape(xx.shape) plt.contourf(xx, yy, Z, alpha=0.2, cmap=plt.cm.coolwarm) # 设置图的标签和标题 plt.xlabel('Dimension 1') plt.ylabel('Dimension 2') plt.title('TF-IDF Features (2D Reduced) & SVM Classification Boundary') plt.legend() plt.grid(True, alpha=0.3) plt.show()
关键细节说明
降维选择:
- PCA是线性降维,计算速度快,适合快速看整体分布,但可能会丢失一些非线性的特征关联
- t-SNE是非线性降维,能更好地保留样本间的局部结构,对文本这种复杂数据的可视化效果通常更好,但计算时间更长(处理2000个样本大概需要几十秒)
超平面绘制的小技巧:
原始的SVM是在高维TF-IDF空间训练的,没法直接映射到2维空间画超平面。所以我们用降维后的2维特征重新训练了一个SVM,这样就能直接在2维网格上生成决策边界——虽然这和原始高维模型不完全一样,但足够帮你直观理解分类逻辑。单独查看单个词的TF-IDF分布:
如果你想看看某个特定词(比如"good")的TF-IDF值分布,可以这样做:# 获取目标词在词汇表中的索引 target_word = "good" word_idx = vectorizer.vocabulary_.get(target_word) if word_idx is not None: # 提取该词的TF-IDF值 word_tfidf = tf_idf_matrix[:, word_idx].toarray().flatten() # 绘制箱线图,对比正负样本的TF-IDF值 plt.boxplot([word_tfidf[labels==0], word_tfidf[labels==1]], labels=['Negative', 'Positive']) plt.title(f'TF-IDF Value of "{target_word}" in Reviews') plt.show() else: print(f"Word '{target_word}' not found in vocabulary!")
内容的提问来源于stack exchange,提问作者OnePunchMan
相关产品推荐
相关产品推荐

