You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark技术求助:从等长字符串数组生成关联行与列

PySpark解决方案

实现步骤

  • 使用arrays_zip将fruits和colors两个数组打包成结构体数组,确保元素一一对应
  • 使用explode将结构体数组展开为多行
  • 使用concat_ws拼接结构体中的水果和颜色字段,生成Connected列
  • 选择最终需要的列

代码示例

from pyspark.sql import SparkSession
from pyspark.sql.functions import arrays_zip, explode, concat_ws

# 初始化SparkSession
spark = SparkSession.builder.appName("ArrayTransform").getOrCreate()

# 创建示例数据集
data = [
    (["banana", "strawberry"], ["yellow", "red"], "good"),
    (["blueberry"], ["blue"], "better"),
    (["melon", "pineapple", "cherry"], ["green", "orange", "red"], "the best")
]
df = spark.createDataFrame(data, ["fruits", "colors", "taste"])

# 执行转换操作
result_df = df \
    .withColumn("fruits_colors", arrays_zip("fruits", "colors")) \
    .withColumn("fruits_colors", explode("fruits_colors")) \
    .withColumn("Connected", concat_ws(" - ", "fruits_colors.fruits", "fruits_colors.colors")) \
    .select("Connected", "taste")

# 显示结果
result_df.show(truncate=False)

代码说明

  • arrays_zip("fruits", "colors"):将两个数组的对应元素打包成结构体,比如[banana, strawberry]和[yellow, red]会变成[{"fruits":"banana","colors":"yellow"}, {"fruits":"strawberry","colors":"red"}]
  • explode("fruits_colors"):把结构体数组拆分成多行,每个结构体对应一行
  • concat_ws(" - ", ...):用指定分隔符拼接两个字段,生成目标字符串列

执行上述代码后,输出结果将与你提供的转换后数据集一致。

内容的提问来源于stack exchange,提问作者MarkusH

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 10:32:05