You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java Spark DataFrame多列去重合并为单列的优化实现

Efficient Java Spark Solution Without UDFs

Great question! Here's a clean, efficient approach using native Java Spark APIs (no UDFs required) that aligns with your requirements:

Core Approach

Instead of handling each column separately and then merging, we can:

  1. Extract each target column (Column_2, Column_3, and filtered Column_5) into a single standardized column
  2. Union these datasets together
  3. Perform a single global deduplication to get the final result

This leverages Spark's built-in optimizations and keeps the code concise.

Java Code Implementation

import org.apache.spark.sql.Dataset;
import org.apache.spark.sql.Row;
import static org.apache.spark.sql.functions.col;

// Assume your original DataFrame is named `sourceDF`
// Step 1: Extract Column_2 as the result column
Dataset<Row> col2Data = sourceDF.select(col("Column_2").alias("result_col"));

// Step 2: Extract Column_3 as the result column
Dataset<Row> col3Data = sourceDF.select(col("Column_3").alias("result_col"));

// Step 3: Extract Column_5 only where Column_4 is 'Y'
Dataset<Row> col5FilteredData = sourceDF
    .filter(col("Column_4").equalTo("Y"))
    .select(col("Column_5").alias("result_col"));

// Step 4: Union all datasets and deduplicate
Dataset<Row> finalResult = col2Data
    .unionByName(col3Data)
    .unionByName(col5FilteredData)
    .distinct();

// View the result
finalResult.show(false);

Key Details

  • unionByName instead of union: This method safely merges DataFrames by column name (not position), making your code more robust if column order ever changes in the source data.
  • Single distinct() call: Instead of deduplicating each column separately, we deduplicate once after unioning. Spark's optimizer will handle this efficiently, and it's cleaner than multiple distinct operations.
  • No UDFs: All logic uses Spark's native column operations and filters, which are better optimized than custom UDFs.

Result Verification

For your sample input data, this code will produce the following single-column DataFrame (order may vary, but values match your expectation):

+----------+
|result_col|
+----------+
|A         |
|B         |
|F         |
|D         |
|C         |
|G         |
|H         |
+----------+

内容的提问来源于stack exchange,提问作者Padfoot123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 13:57:33