Scala Spark DataFrame:保留props结构并扁平化其内部嵌套
props While Keeping props Intact Alright, let's tackle this problem exactly as you described. You need to recursively flatten all nested levels inside the props struct, but retain props as a top-level struct (instead of exploding it into separate columns like common flattening solutions do). This approach works no matter how deep the nesting inside props goes, and it’s fully dynamic so it handles unknown key counts automatically.
Scala Implementation
First, we’ll create a recursive function to traverse the nested structs inside props, generate dot-separated column names for each leaf field, then repackage those fields back into a new props struct.
import org.apache.spark.sql.functions._ import org.apache.spark.sql.types._ // Recursive function to flatten nested structs within a parent column def flattenStructWithin(parentCol: String, structType: StructType): Array[(String, Column)] = { structType.flatMap { field => field.dataType match { case nestedStruct: StructType => // Recursively process nested structs, then strip the parent prefix from column names flattenStructWithin(s"$parentCol.${field.name}", nestedStruct).map { case (colName, colExpr) => (colName.replace(s"$parentCol.", ""), colExpr) } case _ => // For non-struct fields, return the field name and its column reference Array((s"${field.name}", col(s"$parentCol.${field.name}"))) } } } // Sample DataFrame matching your input val df = spark.read.json(Seq( """{"id": "abchchd", "test_id": "ndsbsb", "props": {"type": {"isMale": true, "id": "dd", "mcc": 1234, "name": "Adam"}}}""", """{"id": "abc", "test_id": "asf", "props": {"type2": {"isMale": true, "id": "dd", "mcc": 12134, "name": "Perth"}}}""" ).toDS()) // Extract the struct type of the `props` column val propsStruct = df.schema("props").dataType.asInstanceOf[StructType] // Generate flattened columns for the inner props fields val flattenedPropsCols = flattenStructWithin("props", propsStruct) .map { case (colName, colExpr) => colExpr.as(colName) } // Build the final DataFrame: keep id/test_id, and repackage flattened fields into props val resultDF = df.select( col("id"), col("test_id"), struct(flattenedPropsCols: _*).as("props") ) // Verify the schema and output resultDF.printSchema() resultDF.show(false)
Python Implementation
If you’re working with PySpark, here’s the equivalent recursive solution:
from pyspark.sql import functions as F from pyspark.sql.types import StructType def flatten_struct_within(parent_col, struct_type): fields = [] for field in struct_type.fields: if isinstance(field.dataType, StructType): # Recursively process nested structs nested_fields = flatten_struct_within(f"{parent_col}.{field.name}", field.dataType) # Remove the parent column prefix from field names fields.extend([(col_name.replace(f"{parent_col}.", ""), F.col(col_expr)) for col_name, col_expr in nested_fields]) else: # Handle non-struct leaf fields col_expr = f"{parent_col}.{field.name}" fields.append((field.name, F.col(col_expr))) return fields # Create sample DataFrame data = [ {"id": "abchchd", "test_id": "ndsbsb", "props": {"type": {"isMale": True, "id": "dd", "mcc": 1234, "name": "Adam"}}}, {"id": "abc", "test_id": "asf", "props": {"type2": {"isMale": True, "id": "dd", "mcc": 12134, "name": "Perth"}}} ] df = spark.createDataFrame(data) # Get the struct schema of `props` props_struct = df.schema["props"].dataType # Generate flattened columns for inner props fields flattened_props_cols = flatten_struct_within("props", props_struct) struct_fields = [F.col(col_expr).alias(col_name) for col_name, col_expr in flattened_props_cols] # Assemble the final DataFrame result_df = df.select( F.col("id"), F.col("test_id"), F.struct(*struct_fields).alias("props") ) # Check results result_df.printSchema() result_df.show(truncate=False)
What This Does
- Recursive Flattening: The function traverses every nested struct inside
props, creating dot-separated column names liketype.isMaleortype2.mccfor each leaf field. - Preserve
propsStruct: Instead of moving these flattened fields to the root schema, we repackage them into a newpropsstruct, keeping your desired schema structure. - Dynamic Handling: Works with any number of keys or nesting depth inside
props—no hardcoding required.
Sample Output Schema
The resulting schema will match exactly what you requested:
root |-- id: string (nullable = true) |-- test_id: string (nullable = true) |-- props: struct (nullable = true) | |-- type.id: string (nullable = true) | |-- type.isMale: boolean (nullable = true) | |-- type.mcc: long (nullable = true) | |-- type.name: string (nullable = true) | |-- type2.id: string (nullable = true) | |-- type2.isMale: boolean (nullable = true) | |-- type2.mcc: long (nullable = true) | |-- type2.name: string (nullable = true)
内容的提问来源于stack exchange,提问作者mythic

