PySpark DataFrame dtypes完整列表及数值/类别列区分问询
嘿,这个问题问到点子上了——很多刚上手PySpark做数据预处理的同学都会在类型判断这里卡壳!我来给你捋清楚:
首先:Spark DataFrame
dtypes 返回的完整类型字符串列表 当你调用df.dtypes时,返回的是列名和对应类型字符串的元组列表,所有可能的类型字符串包括这些:
数值类型(和连续型强相关)
'byte'(对应ByteType,1字节整数)'short'(ShortType,2字节整数)'int'(IntegerType,4字节整数)'bigint'(LongType,8字节整数)'float'(FloatType,单精度浮点数)'double'(DoubleType,双精度浮点数)'decimal(p,s)'(DecimalType,定点小数,返回字符串会带具体精度p和刻度s,比如decimal(18,2))
常见类别/离散类型
'string'(StringType)'boolean'(BooleanType,只有True/False两个值,通常归为类别)
日期时间类型
'date'(DateType)'timestamp'(TimestampType)
复杂类型
'array<element_type>'(ArrayType,比如array<string>)'map<key_type,value_type>'(MapType,比如map<string,int>)'struct<field1:type1,field2:type2,...>'(StructType,比如struct<name:string,age:int>)
针对你的需求:优化连续/类别列的判断逻辑
你之前定义的continuous_types确实不全,因为Spark支持多种数值类型。这里给你两种更靠谱的方案:
方案1:基于字符串匹配(简单直接)
把所有数值类型都纳入连续型列表,同时处理DecimalType的特殊格式:
# 覆盖所有数值类型的字符串前缀 continuous_type_prefixes = ('byte', 'short', 'int', 'bigint', 'float', 'double', 'decimal') categorical_types = ('string', 'boolean') def classify_cols_by_dtypes(df): continuous_cols = [] categorical_cols = [] other_cols = [] for col, dtype_str in df.dtypes: # 判断是否为数值类型:要么完全匹配前缀,要么是decimal开头 if any(dtype_str.startswith(prefix) for prefix in continuous_type_prefixes): continuous_cols.append(col) elif dtype_str in categorical_types: categorical_cols.append(col) else: other_cols.append(col) return continuous_cols, categorical_cols, other_cols
方案2:基于Spark DataType类型检查(更严谨)
直接利用Spark的类型体系进行判断,避免字符串匹配的潜在误差:
from pyspark.sql.types import NumericType, StringType, BooleanType def classify_cols_by_type(df): continuous_cols = [] categorical_cols = [] other_cols = [] for col_name in df.columns: dtype = df.schema[col_name].dataType # NumericType是所有数值类型的父类,自动覆盖byte/short/int/bigint/float/double/decimal if isinstance(dtype, NumericType): continuous_cols.append(col_name) # 字符串和布尔值归为类别,可根据业务调整(比如日期是否加入) elif isinstance(dtype, (StringType, BooleanType)): categorical_cols.append(col_name) else: other_cols.append(col_name) return continuous_cols, categorical_cols, other_cols
额外提醒:日期时间类型的归类
Date和Timestamp类型的归属要看你的业务场景:
- 如果是作为时间序列的连续变量(比如分析时间趋势),可以归为连续型;
- 如果是用来分组的离散维度(比如按日期天做聚合),可以归为类别型。
内容的提问来源于stack exchange,提问作者Clock Slave
相关产品推荐
相关产品推荐

