You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Scikit-learn可视化词袋模型下两类文本的决策边界

Got it, let's work through this step by step. The key issue here is that your SVC is trained on the high-dimensional text feature space (after CountVectorizer + TFIDF), but you're visualizing in the 2D PCA-projected space. You can't just train a new SVC on the 2D data (like the iris example) because that wouldn't reflect the actual classifier you're using. Instead, we need to map the decision boundary from the high-dimensional space down to your 2D plot.

Here's the complete, modified code that adds the decision boundary to your existing scatter plot:

from sklearn.datasets import fetch_20newsgroups
from sklearn.feature_extraction.text import CountVectorizer, TfidfTransformer
from sklearn.decomposition import PCA
from sklearn.pipeline import Pipeline
from sklearn.svm import SVC
import matplotlib.pyplot as plt
import numpy as np

# 1. Load data and extract high-dimensional text features (your existing code)
newsgroups_train = fetch_20newsgroups(subset='train', categories=['alt.atheism', 'sci.space'])
pipeline = Pipeline([
    ('vect', CountVectorizer()),
    ('tfidf', TfidfTransformer()),
])
X_highdim = pipeline.fit_transform(newsgroups_train.data).todense()

# 2. Train your SVC on the original high-dimensional features
clf = SVC(random_state=241, kernel='linear')
clf.fit(X_highdim, newsgroups_train.target)

# 3. Project data to 2D with PCA for visualization (your existing code)
pca = PCA(n_components=2).fit(X_highdim)
data2D = pca.transform(X_highdim)

# 4. Create a 2D grid to map the decision boundary
# Extend the plot range slightly to ensure the boundary is fully visible
x_min, x_max = data2D[:, 0].min() - 1, data2D[:, 0].max() + 1
y_min, y_max = data2D[:, 1].min() - 1, data2D[:, 1].max() + 1
# Generate dense grid points (adjust step size for smoother/faster rendering)
xx, yy = np.meshgrid(np.arange(x_min, x_max, 0.1),
                     np.arange(y_min, y_max, 0.1))

# 5. Convert grid points back to the high-dimensional feature space
# We need this because our SVC was trained on high-dimensional data
grid_points_2D = np.c_[xx.ravel(), yy.ravel()]
grid_points_highdim = pca.inverse_transform(grid_points_2D)

# 6. Predict class for each grid point using your trained SVC
Z = clf.predict(grid_points_highdim)
# Reshape predictions to match the grid shape
Z = Z.reshape(xx.shape)

# 7. Plot everything together
plt.figure(figsize=(10, 6))
# Fill regions with class colors (adjust alpha for transparency)
plt.contourf(xx, yy, Z, alpha=0.3, cmap=plt.cm.Paired)
# Overlay the original data points
plt.scatter(data2D[:, 0], data2D[:, 1], c=newsgroups_train.target, edgecolors='k', cmap=plt.cm.Paired)
plt.xlabel('PCA Component 1')
plt.ylabel('PCA Component 2')
plt.title('SVC Decision Boundary on 2D PCA-projected Text Data')
plt.show()

Key Explanations:

  • Inverse PCA Transform: We can't predict directly on the 2D points because our SVC wasn't trained on them. By reversing the PCA transformation, we map the 2D grid back to the high-dimensional text feature space where the classifier operates.
  • Grid Generation: meshgrid creates a dense set of points covering your plot area, so the decision boundary will appear smooth and continuous.
  • Contour Plot: contourf fills areas with different colors based on predicted class, making the boundary clear. You can use contour instead if you only want to draw the boundary line.

Quick Tip:

For better feature quality, you might want to add stop_words='english' to your CountVectorizer to filter out common words that don't add classification value.

内容的提问来源于stack exchange,提问作者Alexandr Bazarov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:32:22