You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现100属性SVM模型2D可视化及多类别SVM特征降维咨询

Visualizing Your Multi-Class SVM with 100+ TF-IDF Features

Absolutely! You can absolutely visualize your multi-class SVM model built on TF-IDF text features—feature dimensionality reduction is exactly the right approach here. Let’s walk through how to do this in Python, with practical code examples.

Core Idea

Your 100+ dimensional TF-IDF features can’t be plotted directly in 2D space. We’ll use dimensionality reduction techniques to project these high-dimensional features down to 2 dimensions, then plot the samples (colored by their class labels) and even visualize the SVM’s decision boundaries if needed.

Step-by-Step Implementation

1. Import Required Libraries

We’ll use scikit-learn for the ML pipeline and matplotlib for plotting:

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.svm import SVC
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE
import matplotlib.pyplot as plt
import numpy as np

2. Prepare Your Data & Generate TF-IDF Features

Replace the placeholder data with your actual text corpus and labels:

# Replace with your real text data and 10 class labels
texts = ["customer support query", "product review positive", "technical documentation", ...]
labels = [0, 1, 2, ..., 9]  # Your 10 distinct class labels

# Generate TF-IDF features (match your 100+ feature setup)
tfidf = TfidfVectorizer(max_features=100)  # Adjust max_features to match your model
X = tfidf.fit_transform(texts).toarray()
y = np.array(labels)

3. Train Your Multi-Class SVM

First, train your original SVM model (just like you already have):

# Train multi-class SVM (using One-vs-One for multi-class handling)
svm_model = SVC(kernel='rbf', decision_function_shape='ovo', random_state=42)
svm_model.fit(X, y)

4. Reduce Dimensions to 2D

We have two great options here—choose based on your needs:

  • PCA: Fast, works well for capturing linear relationships between features.
  • t-SNE: Slower but better at preserving non-linear cluster structures (ideal for text data).
# Option 1: PCA (quick linear reduction)
pca = PCA(n_components=2, random_state=42)
X_2d_pca = pca.fit_transform(X)

# Option 2: t-SNE (better for text clustering visualization)
tsne = TSNE(n_components=2, perplexity=30, random_state=42)
X_2d_tsne = tsne.fit_transform(X)

Note: For large datasets, speed up t-SNE by first using PCA to reduce features to ~50 dimensions, then applying t-SNE.

5. Plot the 2D Projection

Visualize the reduced features, coloring each sample by its class label:

plt.figure(figsize=(12, 6))

# Plot PCA results
plt.subplot(1, 2, 1)
scatter_pca = plt.scatter(X_2d_pca[:, 0], X_2d_pca[:, 1], c=y, cmap='tab10', alpha=0.7)
plt.title("TF-IDF Features (PCA 2D Projection)")
plt.legend(handles=scatter_pca.legend_elements()[0], labels=[str(i) for i in range(10)])

# Plot t-SNE results
plt.subplot(1, 2, 2)
scatter_tsne = plt.scatter(X_2d_tsne[:, 0], X_2d_tsne[:, 1], c=y, cmap='tab10', alpha=0.7)
plt.title("TF-IDF Features (t-SNE 2D Projection)")
plt.legend(handles=scatter_tsne.legend_elements()[0], labels=[str(i) for i in range(10)])

plt.tight_layout()
plt.show()

6. Optional: Visualize SVM Decision Boundaries

To see how your SVM separates classes in the 2D space, train a lightweight SVM on the reduced features and plot its decision boundary:

# Train SVM on t-SNE reduced features
svm_2d = SVC(kernel='rbf', decision_function_shape='ovo', random_state=42)
svm_2d.fit(X_2d_tsne, y)

# Create mesh grid for boundary plotting
h = 0.02  # Step size for mesh
x_min, x_max = X_2d_tsne[:, 0].min() - 1, X_2d_tsne[:, 0].max() + 1
y_min, y_max = X_2d_tsne[:, 1].min() - 1, X_2d_tsne[:, 1].max() + 1
xx, yy = np.meshgrid(np.arange(x_min, x_max, h), np.arange(y_min, y_max, h))

# Predict class for every point in the mesh
Z = svm_2d.predict(np.c_[xx.ravel(), yy.ravel()])
Z = Z.reshape(xx.shape)

# Plot decision boundary + samples
plt.figure(figsize=(8, 6))
plt.contourf(xx, yy, Z, alpha=0.3, cmap='tab10')
plt.scatter(X_2d_tsne[:, 0], X_2d_tsne[:, 1], c=y, cmap='tab10', alpha=0.7)
plt.title("SVM Decision Boundary (t-SNE 2D Space)")
plt.legend(handles=scatter_tsne.legend_elements()[0], labels=[str(i) for i in range(10)])
plt.show()

Key Notes

  • Dimensionality reduction does lose some feature information, but t-SNE is great for preserving the cluster structure of your text classes, which is what you care about for visualization.
  • If your dataset is very large, consider subsampling a portion of your data for faster t-SNE computation.

内容的提问来源于stack exchange,提问作者Madhuri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:17:51