如何使用Scala的foldLeft函数计算TMAX最高的最热日期
Great question! Let's break down how to use foldLeft to find your hottest day, since you already have maxBy and reduceLeft working perfectly.
First, let's clear up a small issue in your existing code: when parsing the CSV, your index for the year field is off. Looking at the header:
Day,JD,Month,Year,PRCP(in),SNOW(in),TAVE (F),TMAX (F),TMIN (F)
The Year column is at index 3 (not 4), so your TempData constructor needs to use p(3).toInt for the year. That'll ensure your parsed data is accurate.
Now, onto foldLeft: unlike reduceLeft which uses the first element of the collection as the starting point, foldLeft lets you define your own initial accumulator value. This makes it more flexible—especially if you need to handle empty collections gracefully.
Option 1: Simple implementation (for non-empty data)
If you're sure your data array will never be empty, you can use the first element as the initial accumulator:
val HottestDay3 = data.foldLeft(data.head) { (currentHottest, nextDay) => // Compare the current hottest day's tmax with the next day's tmax if (nextDay.tmax > currentHottest.tmax) nextDay else currentHottest }
Here's how it works:
- Start with
data.headas the initial "hottest day" - For each subsequent
nextDayin the array, compare itstmaxto thecurrentHottest'stmax - Keep the day with the higher temperature as the new accumulator
- At the end, the accumulator holds the day with the highest
tmax
Option 2: Safe implementation (handles empty data)
If there's a chance your dataset could be empty, use Option[TempData] as the accumulator to avoid runtime errors:
val HottestDay3 = data.foldLeft[Option[TempData]](None) { (acc, day) => acc match { // If we haven't found any days yet, use the current day case None => Some(day) // If we already have a hottest day, compare and keep the warmer one case Some(currentHottest) => if (day.tmax > currentHottest.tmax) Some(day) else acc } } // Then you can print it safely: HottestDay3 match { case Some(day) => println(s"Hot day 3 is $day") case None => println("No weather data available!") }
This approach is more robust—if the dataset is empty, it returns None instead of throwing a NoSuchElementException (which would happen if you tried to use data.head on an empty array).
Putting it all together (corrected full code)
Here's your code with the index fix and the safe foldLeft implementation:
import scala.io.Source case class TempData( day: Int, dayOfYear: Int, month: Int, year: Int, precip: Double, snow: Double, tave: Double, tmax: Double, tmin: Double ) object TempData { def main(args: Array[String]): Unit = { val source = Source.fromFile("C:///DataResearch/SparkScala/MN212142_9392.csv.txt") val lines = source.getLines().drop(1) // Skip header val data = lines.flatMap { line => val p = line.split(",") try { // Fixed index for year (p(3) instead of p(4)) Some(TempData( p(0).toInt, p(1).toInt, p(2).toInt, p(3).toInt, p(4).toDouble, p(5).toDouble, p(6).toDouble, p(7).toDouble, p(8).toDouble )) } catch { // Handle any parsing errors to avoid breaking the whole process case e: Exception => println(s"Failed to parse line: $line") None } }.toArray source.close() val hottestDay1 = data.maxBy(_.tmax) println(s"Hot day 1 is $hottestDay1") val hottestDay2 = data.reduceLeft((d1, d2) => if (d1.tmax >= d2.tmax) d1 else d2) println(s"Hot day 2 is $hottestDay2") val hottestDay3 = data.foldLeft[Option[TempData]](None) { (acc, day) => acc match { case None => Some(day) case Some(current) => if (day.tmax > current.tmax) Some(day) else acc } } hottestDay3 match { case Some(day) => println(s"Hot day 3 is $day") case None => println("No valid weather data found.") } } }
内容的提问来源于stack exchange,提问作者Harshit Kakkar

