Scala Spark:如何删除railroadGreaterFile中与railroadInputFile MEMBER_NUM匹配的记录?
Hey Kiran,
Your approach using a left join with a filter on the null check column is totally valid for removing matching records between the two DataFrames! Let me break down what your code is doing clearly, plus share a more concise alternative that leverages Spark's built-in functionality.
Your Current Code Explanation
Let’s walk through each line to confirm it’s working as intended:
val columnrailroadInputFile = railroadInputFile.withColumn("check", lit("check")): You’re adding a dummy "check" column torailroadInputFile—this acts as a marker to identify which records inrailroadGreaterFilehave a match in the input DataFrame.railroadGreaterFile.join(columnrailroadInputFile, Seq("MEMBER_NUM"), "left"): A left join ensures all records fromrailroadGreaterFileare retained, with matching records fromrailroadInputFileappended. Non-matching records will have null values for all columns pulled fromrailroadInputFile(including your "check" column)..filter($"check".isNull): This filters out any records where the "check" column isn’t null—so you’re left only with records inrailroadGreaterFilethat don’t have a matchingMEMBER_NUMinrailroadInputFile..drop($"check"): Finally, you clean up by removing the dummy column since it’s no longer needed.
A More Concise Alternative: Anti Join
Spark has a built-in anti join type that’s purpose-built for this exact use case. It returns all records from the left DataFrame that have no matches in the right DataFrame on the specified join key, no dummy columns required. Here’s how you can rewrite your code:
val railroadGreaterNotInput = railroadGreaterFile.join( railroadInputFile, Seq("MEMBER_NUM"), "anti" )
This achieves the same result as your original code but is more readable and efficient, cutting out the extra steps of adding and dropping the dummy column.
Both methods will correctly remove records from railroadGreaterFile where MEMBER_NUM exists in railroadInputFile. If you ran into any unexpected behavior with your original code (like odd record counts or data mismatches), feel free to share more details—but the core logic here is solid.
内容的提问来源于stack exchange,提问作者Kiran Kumar

