Spark 1.4 Scala API:基于分隔符生成递进拼接列的需求
Got it, let's work through this problem step by step. Since we're using Spark 1.4— which lacks some of the fancier newer functions like array_join or transform—we'll rely on core tools available in that version: split, UDFs, and iterative column creation.
Core Approach
The key idea is straightforward:
- Split the original delimited string into an array of its components using
-as the separator. - For each position from 1 to 12, generate a new column that contains the cumulative concatenation of the first
nelements from the split array (e.g., first 1 element forNew_col1, first 2 forNew_col2, etc.). - Handle cases where the original string has fewer than
ncomponents by returningnullfor those columns.
Full Code Implementation
First, import the necessary Spark SQL functions:
import org.apache.spark.sql.functions._ import org.apache.spark.sql.types._
Next, define a UDF to handle the cumulative concatenation. This UDF takes the split array and the number of elements to include, then returns the joined string (or null if the array is too short):
val getCumulativeString = udf((components: Seq[String], numElements: Int) => { if (components == null || components.length < numElements) { null // Return null if there aren't enough components } else { components.take(numElements).mkString("-") } })
Now, apply this to your DataFrame. Let's assume your input DataFrame is named sourceDF and the target column is original_col:
// First, split the original column into an array val splitComponents = split(sourceDF("original_col"), "-") // Iterate from 1 to 12 to create each cumulative column var resultDF = sourceDF for (i <- 1 to 12) { resultDF = resultDF.withColumn( s"New_col$i", getCumulativeString(splitComponents, lit(i)) ) }
How It Works
- Split Step: The
splitfunction breaks your string like"ABC-DEF-PQR-XYZ"intoArray("ABC", "DEF", "PQR", "XYZ"). - UDF Logic: For each
i(1 to 12), the UDF takes the firstielements of the array and joins them back with-. If the array has only 4 elements (like the example), columnsNew_col5toNew_col12will benull. - Loop: The loop automates creating all 12 columns without writing repetitive code for each one.
Example Output
For an input row with original_col = "ABC-DEF-PQR-XYZ", the output columns will be:
New_col1:"ABC"New_col2:"ABC-DEF"New_col3:"ABC-DEF-PQR"New_col4:"ABC-DEF-PQR-XYZ"New_col5toNew_col12:null
This approach is efficient, compatible with Spark 1.4, and handles all edge cases (strings with 1 to 12 components, null values in the original column).
内容的提问来源于stack exchange,提问作者Rajdip

