You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将PySpark中MapType列的键值对反转并展开重复键?

实现PySpark DataFrame中Map的键值反转

要实现将原始Map中「键→数组值」的结构转换为「数组元素→原键」的Map结构,可以通过以下步骤完成:

解决思路

  1. 拆分Map结构:先将Map拆分为单独的键值对行,每个行包含原键和对应的数组。
  2. 拆分数组元素:再将每个数组拆分为单独的元素行,形成「原键→数组元素」的一对一映射。
  3. 重新聚合为Map:将拆分后的映射关系重新组合成目标Map结构。

完整代码实现

from pyspark.sql import Row
from pyspark.sql.functions import explode, map_from_entries, collect_list, struct, monotonically_increasing_id

# 构造原始DataFrame
df = spark.createDataFrame([
    Row(rules={1: [7408, 4586], 2: [9670, 2474]}),
])

# 步骤1:添加唯一ID(处理多场景下的原始行追踪)
df_with_id = df.withColumn("id", monotonically_increasing_id())

# 步骤2:拆分Map和数组,得到原键与数组元素的一对一映射
df_exploded = df_with_id.select(
    "id",
    explode("rules").alias("original_key", "numbers")
).select(
    "id",
    "original_key",
    explode("numbers").alias("new_key")
)

# 步骤3:按原始行ID聚合,重新构造目标Map
result_df = df_exploded.groupBy("id").agg(
    map_from_entries(collect_list(struct("new_key", "original_key"))).alias("rules")
).drop("id")

# 查看结果
result_df.show(truncate=False)

输出结果

+------------------------------------------------+
|rules                                           |
+------------------------------------------------+
|{7408 -> 1, 4586 -> 1, 9670 -> 2, 2474 -> 2}|
+------------------------------------------------+

注意事项

  • 如果原始DataFrame包含多行数据,添加monotonically_increasing_id()可以确保每行的转换结果独立,不会互相干扰。
  • 若数组中存在重复元素,最终Map中只会保留最后一次出现的映射关系(因为Map的键具有唯一性)。

内容的提问来源于stack exchange,提问作者Swati Singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 14:21:58