You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scala Spark:如何删除railroadGreaterFile中与railroadInputFile MEMBER_NUM匹配的记录?

Hey Kiran,

Your approach using a left join with a filter on the null check column is totally valid for removing matching records between the two DataFrames! Let me break down what your code is doing clearly, plus share a more concise alternative that leverages Spark's built-in functionality.

Your Current Code Explanation

Let’s walk through each line to confirm it’s working as intended:

  • val columnrailroadInputFile = railroadInputFile.withColumn("check", lit("check")): You’re adding a dummy "check" column to railroadInputFile—this acts as a marker to identify which records in railroadGreaterFile have a match in the input DataFrame.
  • railroadGreaterFile.join(columnrailroadInputFile, Seq("MEMBER_NUM"), "left"): A left join ensures all records from railroadGreaterFile are retained, with matching records from railroadInputFile appended. Non-matching records will have null values for all columns pulled from railroadInputFile (including your "check" column).
  • .filter($"check".isNull): This filters out any records where the "check" column isn’t null—so you’re left only with records in railroadGreaterFile that don’t have a matching MEMBER_NUM in railroadInputFile.
  • .drop($"check"): Finally, you clean up by removing the dummy column since it’s no longer needed.

A More Concise Alternative: Anti Join

Spark has a built-in anti join type that’s purpose-built for this exact use case. It returns all records from the left DataFrame that have no matches in the right DataFrame on the specified join key, no dummy columns required. Here’s how you can rewrite your code:

val railroadGreaterNotInput = railroadGreaterFile.join(
  railroadInputFile,
  Seq("MEMBER_NUM"),
  "anti"
)

This achieves the same result as your original code but is more readable and efficient, cutting out the extra steps of adding and dropping the dummy column.

Both methods will correctly remove records from railroadGreaterFile where MEMBER_NUM exists in railroadInputFile. If you ran into any unexpected behavior with your original code (like odd record counts or data mismatches), feel free to share more details—but the core logic here is solid.

内容的提问来源于stack exchange,提问作者Kiran Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:40:17