Spark中如何将DataFrame的yyyyMMdd格式日期转为yyyy-MM-dd?
Great job figuring out the string manipulation for a single value! To apply that logic to your entire DataFrame, you have two solid options—let’s walk through both.
Option 1: Use Spark's Built-in Date Functions (Recommended)
Spark has optimized, native functions for date handling that outperform custom code in most cases (they’re faster and more reliable for large datasets). Here’s how to use them:
First, convert the date column (a string in yyyyMMdd format) to a proper Date type using to_date with the correct format specifier. Then, convert it back to a string in yyyy-MM-dd format using date_format.
import org.apache.spark.sql.functions.{to_date, date_format} // Assume your original DataFrame is named `df` val formattedDf = df.withColumn( "formatted_date", date_format(to_date($"date", "yyyyMMdd"), "yyyy-MM-dd") ) // To replace the original `date` column instead of adding a new one: val updatedDf = df.withColumn( "date", date_format(to_date($"date", "yyyyMMdd"), "yyyy-MM-dd") )
Option 2: Use a Custom UDF (Reuse Your String Logic)
If you want to stick with the string manipulation you already worked out, wrap that logic in a User Defined Function (UDF) and apply it to the column:
import org.apache.spark.sql.functions.udf // Define the UDF that applies your patch operations val formatDateUdf = udf((dateStr: String) => { dateStr.patch(4, "-", 0).patch(7, "-", 0) }) // Apply the UDF to create a new formatted column val formattedDf = df.withColumn( "formatted_date", formatDateUdf($"date") ) // Or replace the original column directly: val updatedDf = df.withColumn( "date", formatDateUdf($"date") )
Just keep in mind: UDFs don’t get the same optimization as Spark’s built-in functions, so Option 1 is better for large datasets.
内容的提问来源于stack exchange,提问作者Rishabh Ojha

