You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从具有指定Schema的DataFrame两列生成数组?

Solution for Generating Arrays from Your Spark DataFrame

Based on your provided schema and sample data, I’ll cover two common scenarios to address the slight discrepancy between the schema definition and how your sample data is displayed:


Scenario 1: data is an array of structs (matches your schema)

Your schema defines data as an array of structs with k (key) and v (value) fields (e.g., each entry looks like {"k": "key1", "v": "value1"}). Here are practical ways to generate arrays from this structure:

Example 1: Create an array of key-value strings

Convert each struct in data into a formatted string like "key: value" and keep them in an array:

from pyspark.sql import functions as F

df = df.withColumn(
    "key_value_array",
    F.transform(
        F.col("data"),
        lambda x: F.concat(x.k, F.lit(": "), x.v)
    )
)

This adds a new column key_value_array with values like ["key1: value1", "key2: value2", ...].

Example 2: Extract keys or values into standalone arrays

Pull just the keys into an array:

df = df.withColumn("keys_array", F.col("data.k"))

Or extract only the values:

df = df.withColumn("values_array", F.col("data.v"))

These create columns like ["key1", "key2", ...] and ["value1", "value2", ...] respectively.

Example 3: Combine multiple columns into a single array

Create an array that includes _id, c, and flattened key-value pairs:

df = df.withColumn(
    "combined_array",
    F.array(
        F.col("_id").cast("string"),
        F.col("c"),
        F.flatten(F.transform(F.col("data"), lambda x: F.array(x.k, x.v)))
    )
)

This produces an array like ["1", "c1", "key1", "value1", "key2", "value2", ...].


Scenario 2: data is an array of two sub-arrays (matches your sample data)

Your sample shows data as [[key1, key2,...], [value1, value2,...]]—an array containing one sub-array of keys and one of values. Here’s how to process this:

Example 1: Convert to struct array (align with your schema)

First, transform the two sub-arrays into the struct format from your schema:

df = df.withColumn(
    "data_structs",
    F.arrays_zip(F.col("data")[0], F.col("data")[1])
).withColumnRenamed("data_structs", "data")

# Rename fields to match your schema's k/v
df = df.withColumn(
    "data",
    F.transform(
        F.col("data"),
        lambda x: F.struct(x["0"].alias("k"), x["1"].alias("v"))
    )
)

Now you can use all the methods from Scenario 1 on this updated data column.

Example 2: Directly create key-value string array

Skip the struct conversion and generate the key-value array directly:

df = df.withColumn(
    "key_value_array",
    F.transform(
        F.arrays_zip(F.col("data")[0], F.col("data")[1]),
        lambda x: F.concat(x["0"], F.lit(": "), x["1"])
    )
)

This gives the same result as Example 1 in Scenario 1.


Quick Notes:

  • Replace df with your actual DataFrame variable name.
  • All operations use PySpark’s built-in functions, optimized for large datasets.

内容的提问来源于stack exchange,提问作者chaouki

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:53:11