Scala程序求助:查找偶数及其前导数的代码无法运行
Hey there! As someone who's fumbled through my share of Scala + Spark mishaps early on, I totally get how stuck you must feel right now—your logic makes sense in your head, but the code won't run, and you can't spot where you went wrong. Let's break down the most common pitfalls here, plus fix the approach to work properly.
First, Let's Cover the Top Mistakes New Scala/Spark Developers Make Here
- Misusing Immutable Lists
Scala's defaultListis immutable. If you tried adding elements with something likeres :: (prevNum, evenNum), that doesn't modify the originalreslist—it returns a brand new list. You either need to reassign the result to your variable, or use a mutable collection likeListBufferinstead. - Ignoring Spark's Distributed Model
Converting an RDD to a List withcollect()works for small datasets, but it pulls all data to your driver node (not ideal for large data). Worse, if you tried modifying a driver-side mutable list inside an RDD transformation (likemap), you'd run into closure issues or unexpected results because Spark executes code on worker nodes. - Off-by-One Index Errors
If your loop starts at index 0 and tries to accesslist(i-1), you'll hit anIndexOutOfBoundsExceptionimmediately. You need to start iterating from index 1 instead.
Let's Fix the Code—Two Approaches
Approach 1: Do It the Spark Way (Distributed, Efficient)
Instead of converting to a List, use Spark's built-in RDD transformations to handle the logic distributedly. Here are two clean methods:
Using sliding() for Windowed Operations
This lets you look at pairs of consecutive elements directly:
// Sample input RDD val numsRDD = sc.parallelize(1 to 10) // Get sliding windows of 2 elements, filter for windows where the second is even val evenWithPredecessorRDD = numsRDD .sliding(2) // Creates RDD[Array[Int]] with pairs like [1,2], [2,3], etc. .filter(window => window(1) % 2 == 0) // Keep only windows where the second element is even .map(window => (window(0), window(1))) // Convert to (predecessor, even number) pairs // View results evenWithPredecessorRDD.collect().foreach(println)
Output will be:
(1,2) (3,4) (5,6) (7,8) (9,10)
Using zipWithIndex for Index-Based Joining
If you need more control over element positions, this works too:
val numsRDD = sc.parallelize(1 to 10) // Attach indices to each element, swap to (index, number) val indexedRDD = numsRDD.zipWithIndex().map(_.swap) // Shift indices by 1 to pair each element with the next one val shiftedRDD = indexedRDD.map { case (idx, num) => (idx + 1, num) } // Join to get (next_index, (predecessor_num, current_num)) val joinedRDD = indexedRDD.join(shiftedRDD) // Filter for even current numbers, then extract the pairs val resultRDD = joinedRDD .filter { case (_, (_, currNum)) => currNum % 2 == 0 } .map { case (_, (prevNum, currNum)) => (prevNum, currNum) } resultRDD.collect().foreach(println)
Approach 2: If You Really Want to Use a List (For Small Data Only)
If you're set on converting the RDD to a List, fix the mutable collection and loop logic:
import scala.collection.mutable.ListBuffer val numsRDD = sc.parallelize(1 to 10) val numsList = numsRDD.collect().toList // Pull data to driver (only for small datasets!) val results = ListBuffer.empty[(Int, Int)] // Start at index 1 to avoid out-of-bounds errors for (i <- 1 until numsList.length) { val currentNum = numsList(i) if (currentNum % 2 == 0) { val predecessor = numsList(i-1) results.append( (predecessor, currentNum) ) // Use append() for mutable ListBuffer } } // Convert back to immutable List if needed val finalResults = results.toList finalResults.foreach(println)
Key Takeaways
- Prefer Spark's distributed transformations over pulling data to the driver (unless your dataset is tiny).
- Remember that Scala's default collections are immutable—if you need to build a collection incrementally, use
ListBufferor reassign variables when using::. - Double-check loop indices to avoid out-of-bounds errors.
内容的提问来源于stack exchange,提问作者Joseph Chacko

