You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark 1.6 DataFrame:将Struct中所有字段值转为大写

Convert All String Fields in Nested Structs (Including Arrays) to Uppercase in Spark 1.6

Got it, dealing with nested structures in Spark 1.6 can be a bit tricky since it lacks the higher-order functions available in later versions, but we can build a recursive solution to handle all string fields in structs—including those inside arrays—and convert them to uppercase. Here's a step-by-step approach:

Core Approach

We'll recursively traverse the DataFrame's schema and apply transformations based on field type:

  • StringType: Use Spark's upper() function to convert values to uppercase.
  • StructType: Recursively process each sub-field of the struct to build a new transformed struct.
  • ArrayType with Struct elements: Use a UDF to iterate over each struct in the array and apply the uppercase transformation to its string fields.
  • Other types (e.g., Integer, Boolean): Leave them unchanged.

Full Code Implementation

First, import the necessary Spark classes and functions:

import org.apache.spark.sql.{DataFrame, Column, Row}
import org.apache.spark.sql.types.{DataType, StructType, ArrayType, StringType}
import org.apache.spark.sql.functions.{col, upper, udf}

Next, define recursive functions to generate transformed columns:

// Generate transformed columns for all root fields of the DataFrame
def getTransformedColumns(df: DataFrame): Array[Column] = {
  df.schema.fields.map(field => transformColumn(field.name, field.dataType))
}

// Recursively transform a column based on its data type
def transformColumn(colName: String, dataType: DataType): Column = {
  dataType match {
    // Handle nested structs: recursively process each sub-field
    case structType: StructType =>
      val transformedStructFields = structType.fields.map(subField => {
        transformColumn(s"$colName.${subField.name}", subField.dataType).as(subField.name)
      })
      struct(transformedStructFields: _*).as(colName)

    // Handle arrays containing structs: use a UDF to process each struct element
    case arrayType: ArrayType if arrayType.elementType.isInstanceOf[StructType] =>
      val elementStructType = arrayType.elementType.asInstanceOf[StructType]
      
      // UDF to convert string fields in a single struct to uppercase (handles nulls)
      val structToUpperUdf = udf((structRow: Row) => {
        val transformedValues = elementStructType.fields.map(field => {
          field.dataType match {
            case StringType =>
              val value = structRow.getAs[String](field.name)
              if (value == null) null else value.toUpperCase()
            case _ => structRow.get(field.name)
          }
        })
        Row.fromSeq(transformedValues)
      }, elementStructType)
      
      // Apply the UDF to every element in the array
      structToUpperUdf(col(colName)).as(colName)

    // Handle string fields: convert to uppercase
    case StringType =>
      upper(col(colName)).as(colName)

    // Leave all other data types unchanged
    case _ =>
      col(colName).as(colName)
  }
}

How to Use It

Apply the transformation to your original DataFrame:

// Assume your original DataFrame is named `rawDF`
val transformedDF = rawDF.select(getTransformedColumns(rawDF): _*)

Key Notes

  • Null Safety: The UDF checks for null string values to avoid NullPointerException—critical since your schema marks fields as nullable.
  • Deep Nested Structures: The recursive logic handles structs nested at any depth (e.g., address.contact.name will also be converted to uppercase).
  • Spark 1.6 Compatibility: This solution avoids using higher-order functions like transform() (introduced in Spark 2.3) which aren't available in 1.6.

内容的提问来源于stack exchange,提问作者SSC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:35:33