You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从PySpark DataFrame生成指定格式的字典列表

解决方案

你需要先按id和rank分组,将同rank下的types(替换下划线为空格)与social_media构建键值对,再补充rank字段,最后按id收集这些字典为列表。具体代码如下:

import pyspark.sql.functions as F

# 替换types中的下划线为空格,并构建每个rank对应的键值对
df_result = df1.groupBy("id", "rank").agg(
    F.map_from_entries(
        F.collect_list(F.struct(F.regexp_replace("types", "_", " "), "social_media"))
    ).alias("rank_map")
).withColumn(
    # 给每个rank_map添加rank字段
    "final_dict",
    F.map_concat(F.create_map(F.lit("rank"), "rank"), "rank_map")
).groupBy("id").agg(
    # 收集每个id下的所有字典为列表
    F.collect_list("final_dict").alias("result_list")
)

df_result.show(truncate=False)

输出结果:

+---+-----------------------------------------------------------------------+
|id |result_list                                                             |
+---+-----------------------------------------------------------------------+
|1  |[{rank -> 1, search engine -> google, social media -> pinterest}, {rank -> 2, search engine -> yahoo, social media -> youtube}]|
+---+-----------------------------------------------------------------------+

步骤说明:

  1. 处理types字段并按id+rank分组:用regexp_replace将types中的下划线替换为空格,通过collect_list+struct收集同rank下的(类型, 平台)键值对,再用map_from_entries转为map结构。
  2. 补充rank字段到字典:使用map_concat将包含rank的新map和之前的rank_map合并,得到符合要求的完整字典。
  3. 按id收集列表:最后按id分组,用collect_list把每个rank对应的字典收集成列表,得到目标格式。

内容的提问来源于stack exchange,提问作者Praveen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 19:43:14