You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用SciKit-Learn可视化高维Support Vector Machine(SVM)边界?

高维SVM边界可视化的可行方案

当数据维度高于2时,无法直接绘制高维决策边界,我们可以通过以下几种方法间接可视化,理解SVM的运行机制:

方法1:固定其他特征,可视化两两特征的局部决策边界

核心思路是:选择一对特征,将其余所有特征固定在均值、中位数或其他有代表性的数值上,然后在这一对特征的二维空间中绘制SVM的决策边界。这种方法能直观展示某两个特征在其他特征固定时的分类边界,适合分析特征间的交互。

代码示例(以dim=10为例)

import numpy as np
import matplotlib.pyplot as plt
from sklearn.svm import SVC
from sklearn.datasets import make_blobs

kernel='rbf'
dim = 10

np.random.seed(0)
X, y = make_blobs(n_samples=40, centers=4, random_state=6, n_features=dim)

# 训练高维SVM
clf = SVC(kernel=kernel, degree=3, C=1, decision_function_shape='ovr', probability=True)
clf.fit(X, y)

# 选择要可视化的两个特征索引(比如前两个)
feat1_idx, feat2_idx = 0, 1
# 计算其余特征的均值,作为固定值
fixed_features = np.mean(X, axis=0)

# 创建绘图用的二维网格
x_min, x_max = X[:, feat1_idx].min() - 1, X[:, feat1_idx].max() + 1
y_min, y_max = X[:, feat2_idx].min() - 1, X[:, feat2_idx].max() + 1
xx, yy = np.meshgrid(np.arange(x_min, x_max, 0.02),
                     np.arange(y_min, y_max, 0.02))

# 构造高维输入:网格点 + 固定的其他特征
grid = np.c_[xx.ravel(), yy.ravel()]
full_grid = np.repeat(fixed_features.reshape(1, -1), grid.shape[0], axis=0)
full_grid[:, feat1_idx] = grid[:, 0]
full_grid[:, feat2_idx] = grid[:, 1]

# 用训练好的SVM预测网格点类别
Z = clf.predict(full_grid)
Z = Z.reshape(xx.shape)

# 绘制边界和样本点
fig, ax = plt.subplots()
ax.contourf(xx, yy, Z, cmap=plt.cm.coolwarm, alpha=0.8)
ax.scatter(X[:, feat1_idx], X[:, feat2_idx], c=y, cmap=plt.cm.coolwarm, s=20, edgecolors="k")
ax.set_xlabel(f"Feature {feat1_idx+1}")
ax.set_ylabel(f"Feature {feat2_idx+1}")
ax.set_title(f"SVM Decision Boundary (Features {feat1_idx+1}&{feat2_idx+1}, Others Fixed at Mean)")
plt.show()

你可以循环遍历多对特征,生成多个子图,全面观察不同特征对的边界情况。

方法2:通过降维算法将高维边界投影到2D空间

使用PCA、t-SNE等降维算法,将高维数据和SVM的决策边界投影到二维空间。需要注意的是,降维后的边界是原高维边界的近似投影,并非严格的高维边界,但能帮助你整体理解SVM的分类效果。

代码示例(PCA降维)

import numpy as np
import matplotlib.pyplot as plt
from sklearn.svm import SVC
from sklearn.datasets import make_blobs
from sklearn.decomposition import PCA

kernel='rbf'
dim = 10

np.random.seed(0)
X, y = make_blobs(n_samples=40, centers=4, random_state=6, n_features=dim)

# 训练高维SVM
clf = SVC(kernel=kernel, degree=3, C=1, decision_function_shape='ovr', probability=True)
clf.fit(X, y)

# 用PCA将高维数据降维到2D
pca = PCA(n_components=2)
X_2d = pca.fit_transform(X)

# 创建2D网格,并通过PCA逆变换还原为高维
x_min, x_max = X_2d[:, 0].min() - 1, X_2d[:, 0].max() + 1
y_min, y_max = X_2d[:, 1].min() - 1, X_2d[:, 1].max() + 1
xx, yy = np.meshgrid(np.arange(x_min, x_max, 0.05),
                     np.arange(y_min, y_max, 0.05))
grid_2d = np.c_[xx.ravel(), yy.ravel()]
grid_highdim = pca.inverse_transform(grid_2d)

# 预测网格点类别
Z = clf.predict(grid_highdim)
Z = Z.reshape(xx.shape)

# 绘制投影后的边界和样本点
fig, ax = plt.subplots()
ax.contourf(xx, yy, Z, cmap=plt.cm.coolwarm, alpha=0.8)
ax.scatter(X_2d[:, 0], X_2d[:, 1], c=y, cmap=plt.cm.coolwarm, s=20, edgecolors="k")
ax.set_xlabel("PCA Component 1")
ax.set_ylabel("PCA Component 2")
ax.set_title(f"SVM Boundary Projected to 2D via PCA")
plt.show()

如果是高维数据(如dim=200),t-SNE可能比PCA更适合展示样本的聚类结构,但t-SNE没有逆变换,无法直接生成高维网格点,这时可以仅可视化样本的降维分布和分类结果:

# 仅可视化降维后的样本分类结果
from sklearn.manifold import TSNE

tsne = TSNE(n_components=2, random_state=0)
X_tsne = tsne.fit_transform(X)

fig, ax = plt.subplots()
scatter = ax.scatter(X_tsne[:, 0], X_tsne[:, 1], c=y, cmap=plt.cm.coolwarm, s=20, edgecolors="k")
# 添加预测类别标签(可选)
y_pred = clf.predict(X)
for i, (x, y_coord) in enumerate(X_tsne):
    ax.text(x+0.1, y_coord+0.1, str(y_pred[i]), fontsize=8)
ax.set_title(f"SVM Classification Results (t-SNE Projection)")
plt.legend(*scatter.legend_elements(), title="Classes")
plt.show()

方法3:可视化SVM的决策函数置信度

SVM的决策函数值代表样本到超平面的距离,我们可以将这个值与降维后的样本坐标结合,用颜色深浅表示置信度,直观展示哪些样本离决策边界更近。

代码示例

import numpy as np
import matplotlib.pyplot as plt
from sklearn.svm import SVC
from sklearn.datasets import make_blobs
from sklearn.decomposition import PCA

kernel='rbf'
dim = 10

np.random.seed(0)
X, y = make_blobs(n_samples=40, centers=4, random_state=6, n_features=dim)

clf = SVC(kernel=kernel, degree=3, C=1, decision_function_shape='ovr', probability=True)
clf.fit(X, y)

# 计算每个样本的决策函数值(多分类下是每个类别的距离)
decision_vals = clf.decision_function(X)
# 取每个样本到预测类别的距离作为置信度
confidence = np.max(decision_vals, axis=1)

# PCA降维
pca = PCA(n_components=2)
X_2d = pca.fit_transform(X)

# 绘制带置信度的散点图
fig, ax = plt.subplots()
scatter = ax.scatter(X_2d[:, 0], X_2d[:, 1], c=confidence, cmap='viridis', s=50, edgecolors="k")
plt.colorbar(scatter, label="Confidence (Distance to Decision Boundary)")
ax.set_title(f"SVM Confidence Distribution (PCA Projection)")
plt.show()

内容的提问来源于stack exchange,提问作者Triceratops

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 15:30:58