如何将RDD转换为POJO的Java列表?基于已有PairRDD的实现方案
Hey there! Since you're already halfway done with the grouping and summing part using Tuple2, let's walk through turning that JavaPairRDD into a clean list of your custom POJOs step by step.
1. First, Define Your POJO Class
You'll need a simple Java class to hold the grouped attributes and the summed value. Let's call it GroupedResult—make sure it has fields matching your key's data, plus constructors, getters, and setters (a toString method is optional but super helpful for debugging):
public class GroupedResult { private Integer col1; private String col2; private Integer sumCol3; // Constructor that matches your RDD's key and value structure public GroupedResult(Integer col1, String col2, Integer sumCol3) { this.col1 = col1; this.col2 = col2; this.sumCol3 = sumCol3; } // Getters and setters (required if you need to access fields later, or for Spark's serialization) public Integer getCol1() { return col1; } public void setCol1(Integer col1) { this.col1 = col1; } public String getCol2() { return col2; } public void setCol2(String col2) { this.col2 = col2; } public Integer getSumCol3() { return sumCol3; } public void setSumCol3(Integer sumCol3) { this.sumCol3 = sumCol3; } // Optional: For easier logging and debugging @Override public String toString() { return "GroupedResult{" + "col1=" + col1 + ", col2='" + col2 + '\'' + ", sumCol3=" + sumCol3 + '}'; } }
2. Map the JavaPairRDD to Your POJO
Take your existing pairRDD and use the map() transformation to convert each key-value pair into an instance of GroupedResult. The key is a Tuple2<Integer, String> (holding col1 and col2), and the value is your summed col3:
JavaRDD<GroupedResult> resultRDD = pairRDD.map(entry -> { // Extract the grouped key values from the Tuple2 Tuple2<Integer, String> groupedKey = entry._1(); // Extract the summed value Integer totalSum = entry._2(); // Return a new POJO instance with all the data return new GroupedResult(groupedKey._1(), groupedKey._2(), totalSum); });
3. Collect the RDD into a Local Java List
Finally, use the collect() action to pull the distributed data from Spark into a local List<GroupedResult>:
List<GroupedResult> pojoList = resultRDD.collect();
Quick Note on Matching Your Original SQL Query
Your current RDD setup groups by col1 and col2 and sums col3, which aligns with a SQL query like:
SELECT col1, col2, SUM(col3) FROM table GROUP BY col1, col2
If you intended to match your original SQL example (SELECT col1, col2, col3, SUM(col4) FROM table GROUP BY col1, col2, col3), you'd need to adjust your initial mapToPair step to use a Tuple3<Integer, String, [Col3DataType]> as the key (holding all three grouped columns), with the value being col4. Then update your GroupedResult POJO to include col3 and sumCol4 fields, and adjust the map step to extract all three values from the Tuple3.
内容的提问来源于stack exchange,提问作者Vijay Kumar Chauhan

