Spark 1.6 DataFrame:将Struct中所有字段值转为大写
Convert All String Fields in Nested Structs (Including Arrays) to Uppercase in Spark 1.6
Got it, dealing with nested structures in Spark 1.6 can be a bit tricky since it lacks the higher-order functions available in later versions, but we can build a recursive solution to handle all string fields in structs—including those inside arrays—and convert them to uppercase. Here's a step-by-step approach:
Core Approach
We'll recursively traverse the DataFrame's schema and apply transformations based on field type:
- StringType: Use Spark's
upper()function to convert values to uppercase. - StructType: Recursively process each sub-field of the struct to build a new transformed struct.
- ArrayType with Struct elements: Use a UDF to iterate over each struct in the array and apply the uppercase transformation to its string fields.
- Other types (e.g., Integer, Boolean): Leave them unchanged.
Full Code Implementation
First, import the necessary Spark classes and functions:
import org.apache.spark.sql.{DataFrame, Column, Row} import org.apache.spark.sql.types.{DataType, StructType, ArrayType, StringType} import org.apache.spark.sql.functions.{col, upper, udf}
Next, define recursive functions to generate transformed columns:
// Generate transformed columns for all root fields of the DataFrame def getTransformedColumns(df: DataFrame): Array[Column] = { df.schema.fields.map(field => transformColumn(field.name, field.dataType)) } // Recursively transform a column based on its data type def transformColumn(colName: String, dataType: DataType): Column = { dataType match { // Handle nested structs: recursively process each sub-field case structType: StructType => val transformedStructFields = structType.fields.map(subField => { transformColumn(s"$colName.${subField.name}", subField.dataType).as(subField.name) }) struct(transformedStructFields: _*).as(colName) // Handle arrays containing structs: use a UDF to process each struct element case arrayType: ArrayType if arrayType.elementType.isInstanceOf[StructType] => val elementStructType = arrayType.elementType.asInstanceOf[StructType] // UDF to convert string fields in a single struct to uppercase (handles nulls) val structToUpperUdf = udf((structRow: Row) => { val transformedValues = elementStructType.fields.map(field => { field.dataType match { case StringType => val value = structRow.getAs[String](field.name) if (value == null) null else value.toUpperCase() case _ => structRow.get(field.name) } }) Row.fromSeq(transformedValues) }, elementStructType) // Apply the UDF to every element in the array structToUpperUdf(col(colName)).as(colName) // Handle string fields: convert to uppercase case StringType => upper(col(colName)).as(colName) // Leave all other data types unchanged case _ => col(colName).as(colName) } }
How to Use It
Apply the transformation to your original DataFrame:
// Assume your original DataFrame is named `rawDF` val transformedDF = rawDF.select(getTransformedColumns(rawDF): _*)
Key Notes
- Null Safety: The UDF checks for null string values to avoid
NullPointerException—critical since your schema marks fields as nullable. - Deep Nested Structures: The recursive logic handles structs nested at any depth (e.g.,
address.contact.namewill also be converted to uppercase). - Spark 1.6 Compatibility: This solution avoids using higher-order functions like
transform()(introduced in Spark 2.3) which aren't available in 1.6.
内容的提问来源于stack exchange,提问作者SSC
相关产品推荐
相关产品推荐

