You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scala Spark DataFrame:保留props结构并扁平化其内部嵌套

Solution to Flatten Nested Structs Inside props While Keeping props Intact

Alright, let's tackle this problem exactly as you described. You need to recursively flatten all nested levels inside the props struct, but retain props as a top-level struct (instead of exploding it into separate columns like common flattening solutions do). This approach works no matter how deep the nesting inside props goes, and it’s fully dynamic so it handles unknown key counts automatically.

Scala Implementation

First, we’ll create a recursive function to traverse the nested structs inside props, generate dot-separated column names for each leaf field, then repackage those fields back into a new props struct.

import org.apache.spark.sql.functions._
import org.apache.spark.sql.types._

// Recursive function to flatten nested structs within a parent column
def flattenStructWithin(parentCol: String, structType: StructType): Array[(String, Column)] = {
  structType.flatMap { field =>
    field.dataType match {
      case nestedStruct: StructType =>
        // Recursively process nested structs, then strip the parent prefix from column names
        flattenStructWithin(s"$parentCol.${field.name}", nestedStruct).map { case (colName, colExpr) =>
          (colName.replace(s"$parentCol.", ""), colExpr)
        }
      case _ =>
        // For non-struct fields, return the field name and its column reference
        Array((s"${field.name}", col(s"$parentCol.${field.name}")))
    }
  }
}

// Sample DataFrame matching your input
val df = spark.read.json(Seq(
  """{"id": "abchchd", "test_id": "ndsbsb", "props": {"type": {"isMale": true, "id": "dd", "mcc": 1234, "name": "Adam"}}}""",
  """{"id": "abc", "test_id": "asf", "props": {"type2": {"isMale": true, "id": "dd", "mcc": 12134, "name": "Perth"}}}"""
).toDS())

// Extract the struct type of the `props` column
val propsStruct = df.schema("props").dataType.asInstanceOf[StructType]

// Generate flattened columns for the inner props fields
val flattenedPropsCols = flattenStructWithin("props", propsStruct)
  .map { case (colName, colExpr) => colExpr.as(colName) }

// Build the final DataFrame: keep id/test_id, and repackage flattened fields into props
val resultDF = df.select(
  col("id"),
  col("test_id"),
  struct(flattenedPropsCols: _*).as("props")
)

// Verify the schema and output
resultDF.printSchema()
resultDF.show(false)

Python Implementation

If you’re working with PySpark, here’s the equivalent recursive solution:

from pyspark.sql import functions as F
from pyspark.sql.types import StructType

def flatten_struct_within(parent_col, struct_type):
    fields = []
    for field in struct_type.fields:
        if isinstance(field.dataType, StructType):
            # Recursively process nested structs
            nested_fields = flatten_struct_within(f"{parent_col}.{field.name}", field.dataType)
            # Remove the parent column prefix from field names
            fields.extend([(col_name.replace(f"{parent_col}.", ""), F.col(col_expr)) for col_name, col_expr in nested_fields])
        else:
            # Handle non-struct leaf fields
            col_expr = f"{parent_col}.{field.name}"
            fields.append((field.name, F.col(col_expr)))
    return fields

# Create sample DataFrame
data = [
    {"id": "abchchd", "test_id": "ndsbsb", "props": {"type": {"isMale": True, "id": "dd", "mcc": 1234, "name": "Adam"}}},
    {"id": "abc", "test_id": "asf", "props": {"type2": {"isMale": True, "id": "dd", "mcc": 12134, "name": "Perth"}}}
]
df = spark.createDataFrame(data)

# Get the struct schema of `props`
props_struct = df.schema["props"].dataType

# Generate flattened columns for inner props fields
flattened_props_cols = flatten_struct_within("props", props_struct)
struct_fields = [F.col(col_expr).alias(col_name) for col_name, col_expr in flattened_props_cols]

# Assemble the final DataFrame
result_df = df.select(
    F.col("id"),
    F.col("test_id"),
    F.struct(*struct_fields).alias("props")
)

# Check results
result_df.printSchema()
result_df.show(truncate=False)

What This Does

  • Recursive Flattening: The function traverses every nested struct inside props, creating dot-separated column names like type.isMale or type2.mcc for each leaf field.
  • Preserve props Struct: Instead of moving these flattened fields to the root schema, we repackage them into a new props struct, keeping your desired schema structure.
  • Dynamic Handling: Works with any number of keys or nesting depth inside props—no hardcoding required.

Sample Output Schema

The resulting schema will match exactly what you requested:

root
 |-- id: string (nullable = true)
 |-- test_id: string (nullable = true)
 |-- props: struct (nullable = true)
 |    |-- type.id: string (nullable = true)
 |    |-- type.isMale: boolean (nullable = true)
 |    |-- type.mcc: long (nullable = true)
 |    |-- type.name: string (nullable = true)
 |    |-- type2.id: string (nullable = true)
 |    |-- type2.isMale: boolean (nullable = true)
 |    |-- type2.mcc: long (nullable = true)
 |    |-- type2.name: string (nullable = true)

内容的提问来源于stack exchange,提问作者mythic

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 09:13:14