You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark DataFrame dtypes完整列表及数值/类别列区分问询

嘿,这个问题问到点子上了——很多刚上手PySpark做数据预处理的同学都会在类型判断这里卡壳!我来给你捋清楚:

首先:Spark DataFrame dtypes 返回的完整类型字符串列表

当你调用df.dtypes时,返回的是列名和对应类型字符串的元组列表,所有可能的类型字符串包括这些:

数值类型(和连续型强相关)

  • 'byte'(对应ByteType,1字节整数)
  • 'short'(ShortType,2字节整数)
  • 'int'(IntegerType,4字节整数)
  • 'bigint'(LongType,8字节整数)
  • 'float'(FloatType,单精度浮点数)
  • 'double'(DoubleType,双精度浮点数)
  • 'decimal(p,s)'(DecimalType,定点小数,返回字符串会带具体精度p和刻度s,比如decimal(18,2))

常见类别/离散类型

  • 'string'(StringType)
  • 'boolean'(BooleanType,只有True/False两个值,通常归为类别)

日期时间类型

  • 'date'(DateType)
  • 'timestamp'(TimestampType)

复杂类型

  • 'array<element_type>'(ArrayType,比如array<string>)
  • 'map<key_type,value_type>'(MapType,比如map<string,int>)
  • 'struct<field1:type1,field2:type2,...>'(StructType,比如struct<name:string,age:int>)
针对你的需求:优化连续/类别列的判断逻辑

你之前定义的continuous_types确实不全,因为Spark支持多种数值类型。这里给你两种更靠谱的方案:

方案1:基于字符串匹配(简单直接)

把所有数值类型都纳入连续型列表,同时处理DecimalType的特殊格式:

# 覆盖所有数值类型的字符串前缀
continuous_type_prefixes = ('byte', 'short', 'int', 'bigint', 'float', 'double', 'decimal')
categorical_types = ('string', 'boolean')

def classify_cols_by_dtypes(df):
    continuous_cols = []
    categorical_cols = []
    other_cols = []
    
    for col, dtype_str in df.dtypes:
        # 判断是否为数值类型:要么完全匹配前缀,要么是decimal开头
        if any(dtype_str.startswith(prefix) for prefix in continuous_type_prefixes):
            continuous_cols.append(col)
        elif dtype_str in categorical_types:
            categorical_cols.append(col)
        else:
            other_cols.append(col)
    return continuous_cols, categorical_cols, other_cols

方案2:基于Spark DataType类型检查(更严谨)

直接利用Spark的类型体系进行判断,避免字符串匹配的潜在误差:

from pyspark.sql.types import NumericType, StringType, BooleanType

def classify_cols_by_type(df):
    continuous_cols = []
    categorical_cols = []
    other_cols = []
    
    for col_name in df.columns:
        dtype = df.schema[col_name].dataType
        # NumericType是所有数值类型的父类,自动覆盖byte/short/int/bigint/float/double/decimal
        if isinstance(dtype, NumericType):
            continuous_cols.append(col_name)
        # 字符串和布尔值归为类别,可根据业务调整(比如日期是否加入)
        elif isinstance(dtype, (StringType, BooleanType)):
            categorical_cols.append(col_name)
        else:
            other_cols.append(col_name)
    return continuous_cols, categorical_cols, other_cols
额外提醒:日期时间类型的归类

Date和Timestamp类型的归属要看你的业务场景:

  • 如果是作为时间序列的连续变量(比如分析时间趋势),可以归为连续型;
  • 如果是用来分组的离散维度(比如按日期天做聚合),可以归为类别型。

内容的提问来源于stack exchange,提问作者Clock Slave

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:40:22