如何从具有指定Schema的DataFrame两列生成数组?
Based on your provided schema and sample data, I’ll cover two common scenarios to address the slight discrepancy between the schema definition and how your sample data is displayed:
Scenario 1: data is an array of structs (matches your schema)
Your schema defines data as an array of structs with k (key) and v (value) fields (e.g., each entry looks like {"k": "key1", "v": "value1"}). Here are practical ways to generate arrays from this structure:
Example 1: Create an array of key-value strings
Convert each struct in data into a formatted string like "key: value" and keep them in an array:
from pyspark.sql import functions as F df = df.withColumn( "key_value_array", F.transform( F.col("data"), lambda x: F.concat(x.k, F.lit(": "), x.v) ) )
This adds a new column key_value_array with values like ["key1: value1", "key2: value2", ...].
Example 2: Extract keys or values into standalone arrays
Pull just the keys into an array:
df = df.withColumn("keys_array", F.col("data.k"))
Or extract only the values:
df = df.withColumn("values_array", F.col("data.v"))
These create columns like ["key1", "key2", ...] and ["value1", "value2", ...] respectively.
Example 3: Combine multiple columns into a single array
Create an array that includes _id, c, and flattened key-value pairs:
df = df.withColumn( "combined_array", F.array( F.col("_id").cast("string"), F.col("c"), F.flatten(F.transform(F.col("data"), lambda x: F.array(x.k, x.v))) ) )
This produces an array like ["1", "c1", "key1", "value1", "key2", "value2", ...].
Scenario 2: data is an array of two sub-arrays (matches your sample data)
Your sample shows data as [[key1, key2,...], [value1, value2,...]]—an array containing one sub-array of keys and one of values. Here’s how to process this:
Example 1: Convert to struct array (align with your schema)
First, transform the two sub-arrays into the struct format from your schema:
df = df.withColumn( "data_structs", F.arrays_zip(F.col("data")[0], F.col("data")[1]) ).withColumnRenamed("data_structs", "data") # Rename fields to match your schema's k/v df = df.withColumn( "data", F.transform( F.col("data"), lambda x: F.struct(x["0"].alias("k"), x["1"].alias("v")) ) )
Now you can use all the methods from Scenario 1 on this updated data column.
Example 2: Directly create key-value string array
Skip the struct conversion and generate the key-value array directly:
df = df.withColumn( "key_value_array", F.transform( F.arrays_zip(F.col("data")[0], F.col("data")[1]), lambda x: F.concat(x["0"], F.lit(": "), x["1"]) ) )
This gives the same result as Example 1 in Scenario 1.
Quick Notes:
- Replace
dfwith your actual DataFrame variable name. - All operations use PySpark’s built-in functions, optimized for large datasets.
内容的提问来源于stack exchange,提问作者chaouki

