You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Scala获取HDFS中两类目录的最新修改文件?

解决HDFS场景下获取最新修改文件的Scala方案

首先明确:java.io是针对本地文件系统的API,无法直接操作HDFS,必须使用Hadoop官方提供的FileSystem API处理HDFS文件操作。以下是两个场景的具体实现:

场景1:含子目录的dir1(获取最新修改目录内的最新文件)

步骤说明

  • 初始化HDFS文件系统客户端
  • 遍历dir1下的所有子目录,筛选出最新修改的目录
  • 遍历该目录下的所有文件,筛选出最新修改的文件

代码实现

import org.apache.hadoop.fs.{FileSystem, Path, FileStatus}
import org.apache.hadoop.conf.Configuration

object HdfsLatestFileFinder {
  def main(args: Array[String]): Unit = {
    val conf = new Configuration()
    // 集群外运行需指定HDFS地址:conf.set("fs.defaultFS", "hdfs://your-nn-host:9000")
    val fs = FileSystem.get(conf)

    val dir1Path = new Path("/path/to/dir1")

    try {
      // 筛选dir1下的子目录并按修改时间倒序排序
      val subDirs = fs.listStatus(dir1Path)
        .filter(_.isDirectory)
        .sortBy(-_.getModificationTime)

      if (subDirs.nonEmpty) {
        val latestDir = subDirs.head
        // 筛选最新目录下的文件并按修改时间倒序排序
        val filesInLatestDir = fs.listStatus(latestDir.getPath)
          .filter(!_.isDirectory)
          .sortBy(-_.getModificationTime)

        if (filesInLatestDir.nonEmpty) {
          val latestFile = filesInLatestDir.head
          println(s"场景1最新文件:${latestFile.getPath},修改时间:${latestFile.getModificationTime}")
        } else {
          println("最新目录下无文件")
        }
      } else {
        println("dir1下无子目录")
      }
    } catch {
      case e: Exception => e.printStackTrace()
    } finally {
      fs.close()
    }
  }
}

场景2:仅含文件的dir2(直接获取最新修改文件)

步骤说明

直接遍历dir2下的所有文件,按修改时间倒序排序后取第一个即可。

代码实现

// 可复用上述HdfsLatestFileFinder对象,在main方法中添加以下逻辑
val dir2Path = new Path("/path/to/dir2")
try {
  val allFiles = fs.listStatus(dir2Path)
    .filter(!_.isDirectory)
    .sortBy(-_.getModificationTime)

  if (allFiles.nonEmpty) {
    val latestFile = allFiles.head
    println(s"场景2最新文件:${latestFile.getPath},修改时间:${latestFile.getModificationTime}")
  } else {
    println("dir2下无文件")
  }
} catch {
  case e: Exception => e.printStackTrace()
}

注意事项

  • 项目需引入正确的Hadoop客户端依赖,SBT配置示例:
libraryDependencies += "org.apache.hadoop" % "hadoop-client" % "3.3.4" // 版本需与集群Hadoop版本匹配
  • 本地IDE运行时,需在Configuration中指定HDFS的NameNode地址
  • 运行程序的用户需具备HDFS对应目录的读取权限

内容的提问来源于stack exchange,提问作者chandu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 00:12:20