如何在PySpark DataFrame中使用current_timestamp()填充null值
问题场景
DataFrame中名为createdtime的列存在部分null值,需要用当前时间戳填充这些空值。
- 硬编码固定时间字符串传入
fillna的方式可以正常运行,但无法动态获取每次运行时的当前时间:
from pyspark.sql.functions import * default_time = '2022-06-28 05:07:29.077' df = df.fillna({'createdtime': default_time})
- 直接将
current_timestamp()作为填充值传入fillna时会抛出TypeError: Column is not iterable错误,错误代码如下:
from pyspark.sql.functions import * default_time = current_timestamp() df = df.fillna({'createdtime': default_time})
报错截图:
报错原因
fillna 方法的value参数仅支持字面量常量(字符串、数字、布尔值等原生类型),不支持传入Spark Column类型的表达式。current_timestamp()返回的是Column对象,不符合参数要求,因此触发类型错误。
正确实现方式
使用支持Column表达式的空值处理逻辑即可,两种常用写法如下:
写法1:when + isNull 判断
通过when判断列值为空时,返回当前时间戳,否则保留原有列值:
from pyspark.sql.functions import col, current_timestamp, when df = df.withColumn( "createdtime", when(col("createdtime").isNull(), current_timestamp()) .otherwise(col("createdtime")) )
写法2:coalesce 函数
coalesce会返回参数列表中第一个非空值,直接将列本身和当前时间戳作为参数传入即可,代码更简洁:
from pyspark.sql.functions import col, current_timestamp, coalesce df = df.withColumn( "createdtime", coalesce(col("createdtime"), current_timestamp()) )
说明:上述两种写法中,
current_timestamp()会在Spark任务实际执行时取值,所有被填充的空值都会使用任务运行时刻的当前时间,符合动态取值的需求。
内容的提问来源于stack exchange,提问作者Rahul Diggi
相关产品推荐
相关产品推荐

