Spark DataFrame withColumn格式化日份:补0后显示列名而非值求助
Spark DataFrame生成补0日期列的问题解决
你当前代码的问题出在"0" + str(df.DayofMonth)这一行——这是Python层面的字符串拼接操作,str(df.DayofMonth)会把Spark的Column对象直接转成字符串(结果是类似Column<'DayofMonth'>的文本),而不是取列的实际值进行拼接,所以最终生成的列显示的是拼接后的列名字符串,而非预期的补0数值。
解决方案1:使用concat + lit + cast函数
利用Spark内置函数处理列值拼接,替代Python层面的字符串操作:
from pyspark.sql.functions import concat, lit, when, cast dff = df.withColumn( "DayofMonthFormatted", when( df.DayofMonth.between(1, 9), concat(lit("0"), cast(df.DayofMonth, "string")) ).otherwise(cast(df.DayofMonth, "string")) ) display(dff['DayofMonth', 'DayofMonthFormatted'])
lit("0"):创建值为"0"的常量列cast(df.DayofMonth, "string"):将数值类型的DayofMonth转为字符串类型,满足拼接要求concat:Spark的字符串拼接函数,用于合并两个列的内容
解决方案2:使用format_string函数(更简洁)
Spark的format_string函数可直接格式化数字,自动补0,无需额外判断:
from pyspark.sql.functions import format_string dff = df.withColumn( "DayofMonthFormatted", format_string("%02d", df.DayofMonth) ) display(dff['DayofMonth', 'DayofMonthFormatted'])
%02d:格式化规则,将数字转为两位格式,不足两位时前面补0(比如1转为"01",10保持"10")
内容的提问来源于stack exchange,提问作者Rob
相关产品推荐
相关产品推荐

