如何用Pandas识别Sklearn Wine数据集的分类/离散与数值特征
Wine数据集的分类/离散特征识别与处理
一、Wine数据集里的分类/离散特征
Sklearn的Wine数据集里,只有target列是分类/离散特征——它的取值是0、1、2,分别对应三种不同的葡萄酒类别。其余所有特征(比如alcohol酒精含量、malic_acid苹果酸含量等)都是连续数值型的理化检测指标,属于数值特征。
二、Pandas自动识别数字型分类特征的问题与解决思路
Pandas通过dtype确实没法区分数字编码的分类特征和普通数值特征,因为它们的类型都是数值型(int/float)。要解决这个问题,有两种实用思路:
1. 基于数据集文档手动标记
Wine数据集的官方说明里明确标注了target是分类变量,直接手动将其转为分类类型即可:
df['target'] = df['target'].astype('category')
2. 通过唯一值占比自动检测
如果遇到未知数据集,可以通过判断特征的唯一值数量与样本总量的比例来推测——分类特征的唯一值通常远少于样本数,且多为整数编码。可以写个简单的检测函数:
def detect_categorical_features(df, threshold=0.05): categorical_cols = [] total_samples = len(df) for col in df.columns: unique_count = df[col].nunique() # 唯一值占比低于阈值,且为整数类型时判定为分类特征 if unique_count / total_samples < threshold and pd.api.types.is_integer_dtype(df[col]): categorical_cols.append(col) return categorical_cols # 对Wine数据集应用检测 categorical_cols = detect_categorical_features(df) print("识别出的分类特征:", categorical_cols) # 输出会是['target']
三、计算分类特征的取值频率
识别出分类特征后,用Pandas的value_counts()就能快速统计取值的计数或占比:
# 统计取值计数 print(df['target'].value_counts()) # 统计取值频率(占比) print(df['target'].value_counts(normalize=True))
四、提取所有数值特征
排除已识别的分类特征,剩下的就是数值特征:
numerical_cols = [col for col in df.columns if col not in categorical_cols] numerical_features = df[numerical_cols]
内容的提问来源于stack exchange,提问作者Alex Woolfe
相关产品推荐
相关产品推荐

