You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

汽车数据集PCA过程中目标类保留的步骤验证与问题排查

问题分析与修正方案

步骤中的核心问题

1. 主成分与目标类合并的索引不匹配

你创建的pca_df2默认使用连续整数索引,但y_train保留了原数据集随机拆分后的原始索引,直接赋值会触发索引对齐逻辑,导致部分位置出现NaN。

2. 未处理数据缺失

Auto数据集的horsepower列存在缺失值,若跳过处理会导致缩放、PCA计算异常,甚至干扰后续类分离效果。

3. 缩放方式的规范性问题

使用scale()函数虽能完成标准化,但更推荐用StandardScaler类,便于后续将相同的缩放规则统一应用到测试集,保证数据处理的一致性。

修正后的完整步骤

import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.decomposition import PCA
from sklearn.preprocessing import StandardScaler

# 1. 加载并预处理数据(处理缺失值)
auto = pd.read_csv('auto.csv')
# 处理horsepower列缺失值,这里用均值填充,也可选择删除缺失行
auto['horsepower'] = auto['horsepower'].fillna(auto['horsepower'].mean())

# 2. 创建目标类
med = np.median(auto["mpg"])
auto["mpg01"] = auto["mpg"].apply(lambda x: 1 if x > med else 0)

# 3. 拆分数据
X = auto[['cylinders','displacement','horsepower','weight','acceleration','year',"origin"]]
y = auto["mpg01"]
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=101, test_size=0.3, shuffle=True)

# 4. 标准化+PCA降维(用StandardScaler更规范)
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)

pca2 = PCA(n_components=2)
X_train_reduced2 = pca2.fit_transform(X_train_scaled)

# 5. 合并主成分与目标类(解决索引不匹配)
# 方法1:创建DataFrame时指定与y_train一致的索引
pca_df2 = pd.DataFrame(X_train_reduced2, columns=["PC1", "PC2"], index=y_train.index)
pca_df2["mpg01"] = y_train

# 方法2:重置y_train的索引后赋值
# pca_df2 = pd.DataFrame(X_train_reduced2, columns=["PC1", "PC2"])
# pca_df2["mpg01"] = y_train.reset_index(drop=True)

类无分离效果的排查建议

  • 检查主成分解释方差:打印pca2.explained_variance_ratio_,若前两个主成分的解释占比总和低于50%,说明这两个维度不足以区分两类,可尝试增加主成分数量。
  • 分析特征与目标的相关性:计算每个特征与mpg01的相关系数,确认是否存在强关联特征:
    print(X.corrwith(y))
    
  • 验证目标类分布:执行y_train.value_counts(),检查两类数据是否平衡,极端不平衡会导致可视化分离效果差。

内容的提问来源于stack exchange,提问作者David

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 02:20:39