PySpark技术求助:从等长字符串数组生成关联行与列
PySpark解决方案
实现步骤
- 使用
arrays_zip将fruits和colors两个数组打包成结构体数组,确保元素一一对应 - 使用
explode将结构体数组展开为多行 - 使用
concat_ws拼接结构体中的水果和颜色字段,生成Connected列 - 选择最终需要的列
代码示例
from pyspark.sql import SparkSession from pyspark.sql.functions import arrays_zip, explode, concat_ws # 初始化SparkSession spark = SparkSession.builder.appName("ArrayTransform").getOrCreate() # 创建示例数据集 data = [ (["banana", "strawberry"], ["yellow", "red"], "good"), (["blueberry"], ["blue"], "better"), (["melon", "pineapple", "cherry"], ["green", "orange", "red"], "the best") ] df = spark.createDataFrame(data, ["fruits", "colors", "taste"]) # 执行转换操作 result_df = df \ .withColumn("fruits_colors", arrays_zip("fruits", "colors")) \ .withColumn("fruits_colors", explode("fruits_colors")) \ .withColumn("Connected", concat_ws(" - ", "fruits_colors.fruits", "fruits_colors.colors")) \ .select("Connected", "taste") # 显示结果 result_df.show(truncate=False)
代码说明
arrays_zip("fruits", "colors"):将两个数组的对应元素打包成结构体,比如[banana, strawberry]和[yellow, red]会变成[{"fruits":"banana","colors":"yellow"}, {"fruits":"strawberry","colors":"red"}]explode("fruits_colors"):把结构体数组拆分成多行,每个结构体对应一行concat_ws(" - ", ...):用指定分隔符拼接两个字段,生成目标字符串列
执行上述代码后,输出结果将与你提供的转换后数据集一致。
内容的提问来源于stack exchange,提问作者MarkusH
相关产品推荐
相关产品推荐

