扫描近2小时文件夹取最新CSV遇空文件夹报ArrayIndexOutOfBoundsException:0求助
解决文件夹无CSV文件时的数组越界异常
问题根源
异常触发在collectAsList.get(0)这一行:当目标小时文件夹里没有符合条件的CSV文件时,经过过滤、排序、取第一条后的数据集是空的,强行访问索引0的元素就会抛出ArrayIndexOutOfBoundsException: 0。
修复方案
- 先检查文件夹是否存在:避免遍历不存在的HDFS路径导致额外错误
- 判断数据集是否为空:只有当存在符合条件的CSV文件时,才提取最新文件加入列表
修改后的代码
import org.apache.hadoop.conf.Configuration import org.apache.hadoop.fs._ import org.apache.spark.sql._ import org.apache.spark.sql.functions._ import scala.language.postfixOps val hdfsConf = new Configuration() val path = "/user/hdfs/test/input" var finalFiles = List[String]() val currentTs = java.time.LocalDateTime.now val hours = 2 val paths = (0 until hours.toInt).map(h => currentTs.minusHours(h)) .map(ts => s"${path}/partition_date=${ts.toLocalDate}/hour=${ts.toString.substring(11, 13)}") .toList val fs = org.apache.hadoop.fs.FileSystem.get(spark.sparkContext.hadoopConfiguration) for (eachfolder <- paths) { val folderPath = new Path(eachfolder) // 先检查文件夹是否存在且是目录 if (fs.exists(folderPath) && fs.isDirectory(folderPath)) { val pathstatus = fs.listStatus(folderPath) val currpathfiles = pathstatus.map(x => Row(x.getPath.toString, x.getModificationTime)) val latestFileDF = spark.sparkContext.parallelize(currpathfiles) .map(row => (row.getString(0), row.getLong(1))) .toDF("FilePath", "ModificationTime") .filter(col("FilePath").like("%.csv%")) .sort($"ModificationTime".desc) .select(col("FilePath")) .limit(1) // 判断数据集是否为空,再提取元素 if (!latestFileDF.isEmpty) { val latestFile = latestFileDF.map(row => row.getString(0)).collectAsList.get(0) finalFiles = latestFile :: finalFiles } } }
关键优化点
- 把FileSystem实例移到循环外,避免重复创建对象
- 增加文件夹存在性检查,跳过不存在的路径
- 用
latestFileDF.isEmpty判断是否有符合条件的文件,避免空数据集访问索引
内容的提问来源于stack exchange,提问作者Rahul Patidar
相关产品推荐
相关产品推荐

