如何用Scikit-learn可视化词袋模型下两类文本的决策边界
Got it, let's work through this step by step. The key issue here is that your SVC is trained on the high-dimensional text feature space (after CountVectorizer + TFIDF), but you're visualizing in the 2D PCA-projected space. You can't just train a new SVC on the 2D data (like the iris example) because that wouldn't reflect the actual classifier you're using. Instead, we need to map the decision boundary from the high-dimensional space down to your 2D plot.
Here's the complete, modified code that adds the decision boundary to your existing scatter plot:
from sklearn.datasets import fetch_20newsgroups from sklearn.feature_extraction.text import CountVectorizer, TfidfTransformer from sklearn.decomposition import PCA from sklearn.pipeline import Pipeline from sklearn.svm import SVC import matplotlib.pyplot as plt import numpy as np # 1. Load data and extract high-dimensional text features (your existing code) newsgroups_train = fetch_20newsgroups(subset='train', categories=['alt.atheism', 'sci.space']) pipeline = Pipeline([ ('vect', CountVectorizer()), ('tfidf', TfidfTransformer()), ]) X_highdim = pipeline.fit_transform(newsgroups_train.data).todense() # 2. Train your SVC on the original high-dimensional features clf = SVC(random_state=241, kernel='linear') clf.fit(X_highdim, newsgroups_train.target) # 3. Project data to 2D with PCA for visualization (your existing code) pca = PCA(n_components=2).fit(X_highdim) data2D = pca.transform(X_highdim) # 4. Create a 2D grid to map the decision boundary # Extend the plot range slightly to ensure the boundary is fully visible x_min, x_max = data2D[:, 0].min() - 1, data2D[:, 0].max() + 1 y_min, y_max = data2D[:, 1].min() - 1, data2D[:, 1].max() + 1 # Generate dense grid points (adjust step size for smoother/faster rendering) xx, yy = np.meshgrid(np.arange(x_min, x_max, 0.1), np.arange(y_min, y_max, 0.1)) # 5. Convert grid points back to the high-dimensional feature space # We need this because our SVC was trained on high-dimensional data grid_points_2D = np.c_[xx.ravel(), yy.ravel()] grid_points_highdim = pca.inverse_transform(grid_points_2D) # 6. Predict class for each grid point using your trained SVC Z = clf.predict(grid_points_highdim) # Reshape predictions to match the grid shape Z = Z.reshape(xx.shape) # 7. Plot everything together plt.figure(figsize=(10, 6)) # Fill regions with class colors (adjust alpha for transparency) plt.contourf(xx, yy, Z, alpha=0.3, cmap=plt.cm.Paired) # Overlay the original data points plt.scatter(data2D[:, 0], data2D[:, 1], c=newsgroups_train.target, edgecolors='k', cmap=plt.cm.Paired) plt.xlabel('PCA Component 1') plt.ylabel('PCA Component 2') plt.title('SVC Decision Boundary on 2D PCA-projected Text Data') plt.show()
Key Explanations:
- Inverse PCA Transform: We can't predict directly on the 2D points because our SVC wasn't trained on them. By reversing the PCA transformation, we map the 2D grid back to the high-dimensional text feature space where the classifier operates.
- Grid Generation:
meshgridcreates a dense set of points covering your plot area, so the decision boundary will appear smooth and continuous. - Contour Plot:
contourffills areas with different colors based on predicted class, making the boundary clear. You can usecontourinstead if you only want to draw the boundary line.
Quick Tip:
For better feature quality, you might want to add stop_words='english' to your CountVectorizer to filter out common words that don't add classification value.
内容的提问来源于stack exchange,提问作者Alexandr Bazarov

