PySpark如何统计timestamp类型列的null/NaN/空值数量
报错根因
isnan()是专门用来判断浮点型数值是否为NaN的函数,仅支持Float、Double类数值类型字段。你要统计的createdtime是timestamp时间戳类型,本身不存在NaN值存储逻辑,对该类型字段直接调用isnan()会触发类型不匹配报错。
另外timestamp类型的字段不会存储空字符串、NaN这类异常值,无效值只有null,不需要做isnan判断。
可直接运行的正确代码
from pyspark.sql.functions import col, when, count df_columns = ["createdtime"] df.select([ count(when(col(c).isNull(), c)).alias(c) for c in df_columns ]).show()
多类型列兼容写法
如果你后续需要同时统计数值列、时间戳列、字符串列的异常值,可以先判断字段类型再匹配判断逻辑,避免类型报错:
from pyspark.sql.functions import col, isnan, when, count from pyspark.sql.types import FloatType, DoubleType, StringType df_columns = ["createdtime"] def get_invalid_count(col_name): col_type = df.schema[col_name].dataType # 浮点数值列需要判断NaN+null if isinstance(col_type, (FloatType, DoubleType)): judge_cond = isnan(col_name) | col(col_name).isNull() # 字符串列需要判断空串+null elif isinstance(col_type, StringType): judge_cond = (col(col_name).isNull()) | (col(col_name) == "") # 时间戳、整数等其他类型仅判断null else: judge_cond = col(col_name).isNull() return count(when(judge_cond, col_name)).alias(col_name) df.select([get_invalid_count(c) for c in df_columns]).show()
内容的提问来源于stack exchange,提问作者Rahul Diggi
相关产品推荐
相关产品推荐

