Spark-XML:基于Dataframe嵌套数组生成指定结构XML
Hey there! Let's walk through exactly how to generate that nested XML structure using Spark-XML with DataFrames. I'll break this down into clear, actionable steps with code examples in both Python and Scala, since both are common for Spark work.
First, make sure you have the Spark-XML library included in your Spark session. The version needs to match your Spark version—for example, Spark 3.3.x uses spark-xml_2.12:0.15.0, Spark 3.4.x uses 0.16.0.
For Python (pyspark):
Start your session with the package:from pyspark.sql import SparkSession spark = SparkSession.builder \ .appName("NestedXMLGenerator") \ .config("spark.jars.packages", "com.databricks:spark-xml_2.12:0.15.0") \ .getOrCreate()For Scala:
import org.apache.spark.sql.SparkSession val spark = SparkSession.builder .appName("NestedXMLGenerator") .config("spark.jars.packages", "com.databricks:spark-xml_2.12:0.15.0") .getOrCreate()
Your target XML has 3 levels: parent → childs (array of child) → grandchilds (array of grandchild). We need to mirror this structure in Spark's DataFrame schema.
Python Approach (Using StructType)
from pyspark.sql.types import StructType, StructField, StringType, ArrayType # Define the innermost level: grandchild grandchild_schema = StructType([ StructField("name", StringType(), nullable=False) ]) # Define child (with optional grandchilds array) child_schema = StructType([ StructField("name", StringType(), nullable=False), StructField("grandchilds", ArrayType(grandchild_schema), nullable=True) ]) # Define the parent structure parent_struct = StructType([ StructField("name", StringType(), nullable=False), StructField("childs", ArrayType(child_schema), nullable=False) ]) # Wrap the parent in a top-level schema to match our XML root final_schema = StructType([StructField("parent", parent_struct, nullable=False)])
Scala Approach (Using Case Classes)
Case classes make nested structures cleaner in Scala:
import spark.implicits._ // Define case classes for each level case class GrandChild(name: String) case class Child(name: String, grandchilds: Option[Array[GrandChild]]) case class Parent(name: String, childs: Array[Child])
Now populate the DataFrame with your desired data, matching the structure we defined.
Python Example
data = [ ( { "name": "parent", "childs": [ {"name": "child1"}, # Child without grandchilds { "name": "child2", "grandchilds": [ {"name": "grand1"}, {"name": "grand2"}, {"name": "grand3"} ] } # Add more child entries here as needed ] } ) ] df = spark.createDataFrame(data, schema=final_schema)
Scala Example
val parentData = Parent( name = "parent", childs = Array( Child("child1", None), // No grandchilds Child( "child2", Some(Array(GrandChild("grand1"), GrandChild("grand2"), GrandChild("grand3"))) ) ) ) val df = Seq(parentData).toDF("parent")
Use Spark-XML's DataFrameWriter to generate the XML. The key here is setting the right rowTag and rootTag to match your desired structure.
Python Writer Code
df.write \ .format("xml") \ .option("rowTag", "parent") # This sets the tag for our single parent row .option("rootTag", "") # We don't need an extra wrapper around the parent .mode("overwrite") \ .save("/path/to/your/xml/output")
Scala Writer Code
df.write .format("xml") .option("rowTag", "parent") .option("rootTag", "") .mode("overwrite") .save("/path/to/your/xml/output")
- Optional Fields: Spark-XML automatically omits null/None fields (like the missing
grandchildsforchild1), which matches your desired output. - Schema Order: XML tags are generated in the order of your schema, so make sure fields like
namecome beforechildsin your struct definitions. - Empty Arrays: If you have an empty
grandchildsarray, Spark-XML will generate an empty<grandchilds/>tag by default. If you want to omit it entirely, filter out empty arrays before writing. - Dependency Versions: Double-check that your
spark-xmlversion is compatible with your Spark version to avoid runtime errors.
内容的提问来源于stack exchange,提问作者Punith Raj

