You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

真实标签与预测值的降维可视化方案咨询

降维可视化真实与预测标签对应关系的解决方案

问题背景

我有如下结构的DataFrame:

label    predicted     F1  F2   F3 .... F40
major     minor         2   1   4
major     major         1   0   10
minor     patch         4   3   23
major     patch         2   1   11
minor     minor         0   4   8
patch     major         7   3   30
patch     minor         8   0   1
patch     patch         1   7   11

其中label为样本真实标签,predicted为预测标签,另有约40个特征。需求是将40维特征降为2维,可视化真实标签与预测标签的9种组合对应关系。目前使用PCA时2个主成分无法捕获足够方差,且不知如何将PCA结果与原数据集的标签、预测值整体映射;拆分9个数据集处理并非理想方案,寻求其他降维可视化方法。

可行的降维可视化方法

1. t-SNE(t分布邻域嵌入)

t-SNE擅长捕捉数据的局部结构,对高维数据的聚类可视化效果优于PCA,适合展示不同标签组合的分布差异。

  • 操作逻辑:
    • 提取40维特征作为输入,保留原DataFrame的label和predicted列
    • 用t-SNE将特征降维到2维,生成两个新列(如tsne_1、tsne_2)
    • 将降维结果与原标签、预测值合并,直接用散点图可视化,通过颜色+标记区分9种组合(比如用颜色表示真实标签、标记形状表示预测标签,或直接用label_predicted组合字符串作为分组依据)
  • 代码示例:
from sklearn.manifold import TSNE
import pandas as pd
import matplotlib.pyplot as plt

# 假设原数据存储在df中
features = df.drop(['label', 'predicted'], axis=1)
tsne = TSNE(n_components=2, random_state=42)
tsne_results = tsne.fit_transform(features)

# 合并降维结果到原数据集
df['tsne_1'] = tsne_results[:, 0]
df['tsne_2'] = tsne_results[:, 1]

# 可视化9种标签组合
plt.figure(figsize=(10,8))
# 生成标签组合字符串
combinations = df.apply(lambda x: f"{x['label']}_{x['predicted']}", axis=1)
unique_combs = combinations.unique()
# 分配颜色映射
colors = plt.cm.get_cmap('tab10', len(unique_combs))

for i, comb in enumerate(unique_combs):
    subset = df[combinations == comb]
    plt.scatter(subset['tsne_1'], subset['tsne_2'], label=comb, color=colors(i), alpha=0.7)

plt.legend(bbox_to_anchor=(1.05, 1), loc='upper left')
plt.xlabel('t-SNE Component 1')
plt.ylabel('t-SNE Component 2')
plt.title('True vs Predicted Label Combinations (t-SNE)')
plt.show()

2. UMAP(均匀流形近似与投影)

UMAP与t-SNE类似,但保留全局结构的能力更强,计算速度也更快,适合中等规模数据集。

  • 核心优势:相比t-SNE,能更好地保留不同簇之间的全局关系,同时局部聚类效果优异
  • 操作逻辑与t-SNE完全一致:降维后合并原标签列,按9种组合分组绘图即可
  • 代码示例:
import umap
import pandas as pd
import matplotlib.pyplot as plt

features = df.drop(['label', 'predicted'], axis=1)
umap_embedding = umap.UMAP(n_components=2, random_state=42).fit_transform(features)

# 合并结果到原数据集
df['umap_1'] = umap_embedding[:, 0]
df['umap_2'] = umap_embedding[:, 1]

# 可视化
plt.figure(figsize=(10,8))
combinations = df.apply(lambda x: f"{x['label']}_{x['predicted']}", axis=1)
unique_combs = combinations.unique()
colors = plt.cm.get_cmap('tab10', len(unique_combs))

for i, comb in enumerate(unique_combs):
    subset = df[combinations == comb]
    plt.scatter(subset['umap_1'], subset['umap_2'], label=comb, color=colors(i), alpha=0.7)

plt.legend(bbox_to_anchor=(1.05, 1), loc='upper left')
plt.xlabel('UMAP Component 1')
plt.ylabel('UMAP Component 2')
plt.title('True vs Predicted Label Combinations (UMAP)')
plt.show()

3. 自定义特征映射(结合领域知识)

如果40维特征中有领域内的关键特征,可以尝试:

  • 直接选择2个物理意义明确的特征做散点图,按9种组合标记
  • 对特征做线性组合(如加权求和)得到两个综合维度,再可视化,这种方法解释性更强

PCA的优化思路

如果仍想尝试PCA,可做以下调整:

  • 先用StandardScaler对特征做标准化,确保所有特征在同一尺度,避免方差大的特征主导主成分
  • 查看PCA的方差解释率,若前2个主成分不够,可尝试前3个主成分做3D可视化(2D更直观)
  • 将PCA结果与原标签、预测值合并,无需拆分数据集,直接按组合分组绘图

内容的提问来源于stack exchange,提问作者Brie MerryWeather

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 03:15:19