You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark DataFrame按ID分组与另一DataFrame关联的实现求助

解决Spark DataFrame按ID分组关联并补全所有Index的问题

要实现每个ID包含df2的所有index值,匹配取df1的value、未匹配取0的需求,核心是先构造所有ID与所有Index的全量组合,再通过左关联匹配原数据,最后填充默认值。具体实现步骤如下:

步骤1:提取唯一ID集合

从df1中提取所有不重复的id,生成单独的DataFrame:

val uniqueIds = df1.select("id").distinct()

步骤2:构造全量ID-Index组合

将唯一ID集合与df2做交叉连接(crossJoin),得到每个id对应所有index的完整组合:

val fullCombination = uniqueIds.crossJoin(df2.select("index"))

步骤3:左关联原数据并填充默认值

将全量组合与df1左关联,关联条件是id和index同时匹配,然后用coalesce函数优先取df1的value,无匹配时用0填充:

import org.apache.spark.sql.functions._

val result = fullCombination
  .join(df1, Seq("id", "index"), "left")
  .select(
    col("id"),
    col("index"),
    coalesce(col("value"), lit(0)).alias("value")
  )

验证结果

执行result.show()会输出符合预期的结果:

+---+-----+-----+
| id|index|value|
+---+-----+-----+
|  1|    1|  100|
|  1|    2|    0|
|  1|    3|   20|
|  1|    4|    0|
|  1|    5|    0|
|  2|    1|    0|
|  2|    2|   10|
|  2|    3|    0|
|  2|    4|    0|
|  2|    5|    5|
+---+-----+-----+

内容的提问来源于stack exchange,提问作者user1125803

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 04:25:18