You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark 1.4 Scala API:基于分隔符生成递进拼接列的需求

Solution for Cumulative Split Columns in Spark 1.4 (Scala)

Got it, let's work through this problem step by step. Since we're using Spark 1.4— which lacks some of the fancier newer functions like array_join or transform—we'll rely on core tools available in that version: split, UDFs, and iterative column creation.

Core Approach

The key idea is straightforward:

  1. Split the original delimited string into an array of its components using - as the separator.
  2. For each position from 1 to 12, generate a new column that contains the cumulative concatenation of the first n elements from the split array (e.g., first 1 element for New_col1, first 2 for New_col2, etc.).
  3. Handle cases where the original string has fewer than n components by returning null for those columns.

Full Code Implementation

First, import the necessary Spark SQL functions:

import org.apache.spark.sql.functions._
import org.apache.spark.sql.types._

Next, define a UDF to handle the cumulative concatenation. This UDF takes the split array and the number of elements to include, then returns the joined string (or null if the array is too short):

val getCumulativeString = udf((components: Seq[String], numElements: Int) => {
  if (components == null || components.length < numElements) {
    null // Return null if there aren't enough components
  } else {
    components.take(numElements).mkString("-")
  }
})

Now, apply this to your DataFrame. Let's assume your input DataFrame is named sourceDF and the target column is original_col:

// First, split the original column into an array
val splitComponents = split(sourceDF("original_col"), "-")

// Iterate from 1 to 12 to create each cumulative column
var resultDF = sourceDF
for (i <- 1 to 12) {
  resultDF = resultDF.withColumn(
    s"New_col$i",
    getCumulativeString(splitComponents, lit(i))
  )
}

How It Works

  • Split Step: The split function breaks your string like "ABC-DEF-PQR-XYZ" into Array("ABC", "DEF", "PQR", "XYZ").
  • UDF Logic: For each i (1 to 12), the UDF takes the first i elements of the array and joins them back with -. If the array has only 4 elements (like the example), columns New_col5 to New_col12 will be null.
  • Loop: The loop automates creating all 12 columns without writing repetitive code for each one.

Example Output

For an input row with original_col = "ABC-DEF-PQR-XYZ", the output columns will be:

  • New_col1: "ABC"
  • New_col2: "ABC-DEF"
  • New_col3: "ABC-DEF-PQR"
  • New_col4: "ABC-DEF-PQR-XYZ"
  • New_col5 to New_col12: null

This approach is efficient, compatible with Spark 1.4, and handles all edge cases (strings with 1 to 12 components, null values in the original column).

内容的提问来源于stack exchange,提问作者Rajdip

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:57:59