You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark Scala实现按条件填充首行ID至后续行(新增列对比)

Solution in Spark Scala

Got it, let's tackle this problem step by step. The core idea is to create a grouping logic based on non-zero fn values, then propagate the right id through each group until we hit the next non-zero fn row. Here's a practical implementation:

Step 1: Sort Your Data First

Since our logic depends entirely on the order of rows (based on id), we need to sort the DataFrame first to ensure correct processing order.

import org.apache.spark.sql.functions._
import org.apache.spark.sql.expressions.Window

// Assume your input DataFrame is named `df` with columns `id` (Int) and `fn` (Int)
val sortedDf = df.orderBy("id")

Step 2: Get the Next Row's ID

We use the lead window function to fetch the id of the next row—this will be our new "base" ID whenever we encounter a non-zero fn value.

val withNextId = sortedDf.withColumn(
  "next_id",
  lead(col("id"), 1).over(Window.orderBy("id"))
)

Step 3: Mark Trigger Points for ID Switch

Create a column that only stores the next row's id when fn is non-zero. These are our trigger points where we need to switch the propagated ID.

val withTriggerId = withNextId.withColumn(
  "trigger_new_id",
  when(col("fn") =!= 0, col("next_id")).otherwise(null)
)

Step 4: Propagate the ID Through Each Group

Use the last window function (with ignoreNulls=true) to carry forward the most recent trigger ID. For the initial group (before any non-zero fn), we fall back to the first id in the dataset using coalesce.

val resultDf = withTriggerId.withColumn(
  "new_id",
  coalesce(
    last(col("trigger_new_id"), ignoreNulls = true).over(Window.orderBy("id")),
    first(col("id")).over(Window.orderBy("id"))
  )
).drop("next_id", "trigger_new_id") // Clean up intermediate columns we don't need

Step 5: Inspect the Result

You can now view the output to compare the original id with the generated new_id:

resultDf.show()

Example Output

For input data like this:

idfn
10
20
35
40
50
63
70

The result will look like this:

idfnnew_id
101
201
351
404
504
634
707

Handling Edge Cases

  • First row has non-zero fn: The first row uses its own id, then switches to the next row's id for all subsequent rows.
  • No non-zero fn values: Every row will use the first id in the dataset.
  • Last row has non-zero fn: The final row keeps the current group's id (since there's no next row to switch to).

内容的提问来源于stack exchange,提问作者vrreddy1234

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:18:05