You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark如何统计timestamp类型列的null/NaN/空值数量

报错根因

isnan()是专门用来判断浮点型数值是否为NaN的函数,仅支持Float、Double类数值类型字段。你要统计的createdtime是timestamp时间戳类型,本身不存在NaN值存储逻辑,对该类型字段直接调用isnan()会触发类型不匹配报错。
另外timestamp类型的字段不会存储空字符串、NaN这类异常值,无效值只有null,不需要做isnan判断。

可直接运行的正确代码
from pyspark.sql.functions import col, when, count

df_columns = ["createdtime"]
df.select([
    count(when(col(c).isNull(), c)).alias(c) for c in df_columns
]).show()
多类型列兼容写法

如果你后续需要同时统计数值列、时间戳列、字符串列的异常值,可以先判断字段类型再匹配判断逻辑,避免类型报错:

from pyspark.sql.functions import col, isnan, when, count
from pyspark.sql.types import FloatType, DoubleType, StringType

df_columns = ["createdtime"]

def get_invalid_count(col_name):
    col_type = df.schema[col_name].dataType
    # 浮点数值列需要判断NaN+null
    if isinstance(col_type, (FloatType, DoubleType)):
        judge_cond = isnan(col_name) | col(col_name).isNull()
    # 字符串列需要判断空串+null
    elif isinstance(col_type, StringType):
        judge_cond = (col(col_name).isNull()) | (col(col_name) == "")
    # 时间戳、整数等其他类型仅判断null
    else:
        judge_cond = col(col_name).isNull()
    return count(when(judge_cond, col_name)).alias(col_name)

df.select([get_invalid_count(c) for c in df_columns]).show()

内容的提问来源于stack exchange,提问作者Rahul Diggi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 07:51:26