如何选择分类与数值特征并解决特征拼接时的ValueError问题?
问题分析与解决方案
错误根源
你直接对pandas的Index对象(列名集合)使用+运算符,该运算符在Index类型中是元素级别的加法操作,而非列表式的拼接。当两个Index的长度不同(87和42)时,pandas无法完成广播运算,因此抛出ValueError。
修复方法
方法1:转换为列表后拼接
将列名Index转为普通列表,再用+完成拼接:
categorical_features = df.select_dtypes(include=['object', 'category']).columns.tolist() df = df.dropna(subset=categorical_features) numerical_features = df.select_dtypes(include=['int64', 'float64']).columns.tolist() X = df[categorical_features + numerical_features].copy() y = df['JobSatisfaction'].copy()
方法2:使用Index的append方法拼接
利用pandas Index自带的append方法完成列名集合的拼接:
categorical_features = df.select_dtypes(include=['object', 'category']).columns df = df.dropna(subset=categorical_features) numerical_features = df.select_dtypes(include=['int64', 'float64']).columns combined_features = categorical_features.append(numerical_features) X = df[combined_features].copy() y = df['JobSatisfaction'].copy()
方法3:适配新版本pandas(append弃用情况)
如果你的pandas版本≥2.0,append方法已被弃用,可改用pd.concat拼接索引:
import pandas as pd categorical_features = df.select_dtypes(include=['object', 'category']).columns df = df.dropna(subset=categorical_features) numerical_features = df.select_dtypes(include=['int64', 'float64']).columns combined_features = pd.concat([categorical_features, numerical_features]) X = df[combined_features].copy() y = df['JobSatisfaction'].copy()
内容的提问来源于stack exchange,提问作者Rajesh Adhikari
相关产品推荐
相关产品推荐

