You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python筛选特征实现已知样本与其余样本的无重叠可视化

高维样本分组可视化方案

问题背景

现有数据集synthetic_feature_file包含50000+特征与43个样本,需将指定索引(1、6、7、11、14、15、27)的样本与其余样本用不同颜色区分,且两类数据无重叠。原代码采用随机选择特征的方式,无法保证分组区分度,现改用PCA降维方法实现清晰可视化。

改进思路

针对高维小样本数据,用PCA(主成分分析)将50000+特征降维至2维,保留数据的最大方差信息,确保两组样本在二维空间中尽可能分离,满足可视化需求。

完整代码

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
from sklearn.decomposition import PCA
from sklearn.preprocessing import StandardScaler

# 读取数据集
syn_data = pd.read_csv("synthetic_feature_file")
# 标记目标样本索引
sample_indices = syn_data.index.isin([1, 6, 7, 11, 14, 15, 27])

# 数据标准化(PCA前置必要步骤)
scaler = StandardScaler()
scaled_features = scaler.fit_transform(syn_data)

# PCA降维到2维
pca = PCA(n_components=2)
pca_transformed = pca.fit_transform(scaled_features)
pca_df = pd.DataFrame(pca_transformed, columns=["主成分1", "主成分2"])

# 绘制可视化图
plt.figure(figsize=(8, 6))
# 绘制其余样本
plt.scatter(
    pca_df[~sample_indices]["主成分1"],
    pca_df[~sample_indices]["主成分2"],
    color="#1f77b4",
    label="其余样本",
    alpha=0.7
)
# 绘制目标样本(加大尺寸+边框强化区分)
plt.scatter(
    pca_df[sample_indices]["主成分1"],
    pca_df[sample_indices]["主成分2"],
    color="#ff4b5c",
    label="指定样本",
    s=80,
    edgecolor="black"
)

plt.xlabel(f"主成分1(方差占比:{pca.explained_variance_ratio_[0]:.2%})")
plt.ylabel(f"主成分2(方差占比:{pca.explained_variance_ratio_[1]:.2%})")
plt.title("高维样本PCA降维分组可视化")
plt.legend()
plt.grid(alpha=0.3)
plt.show()

代码说明

  • 数据标准化:PCA对数据尺度敏感,标准化后每个特征均值为0、方差为1,避免方差大的特征主导降维结果
  • PCA降维:将高维特征压缩到2维,通过explained_variance_ratio_可查看每个主成分的方差解释占比,直观了解降维保留的信息
  • 可视化优化:给目标样本设置更大尺寸和黑色边框,进一步强化两组区分度;添加网格提升图表可读性

内容的提问来源于stack exchange,提问作者董珈妤

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 08:33:34