You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何快速精准获取ADLS Gen2中Delta表及文件夹的大小?

解决ADLS Gen2中Delta表及文件夹大小统计的最优方案

问题根源分析

你遇到的大小差异是因为微软文档提供的两种方法仅统计Delta表当前快照的活跃数据文件,不包含Delta表的历史版本文件(如旧commit日志、被删除/覆盖的分区文件),而存储资源管理器或dbutils.fs.ls统计的是目录下所有文件(包括历史文件),因此出现67%的差值。

最优解决方案

一、Delta表大小统计

1. 包含历史版本的完整目录大小(最准确)

利用Spark分布式文件遍历替代单线程的dbutils.fs.ls,借助集群executor并行处理,大幅提升效率:

import org.apache.hadoop.fs._
val deltaPath = "abfss://<container>@<storage-account>.dfs.core.windows.net/<delta-table-path>"
val fs = FileSystem.get(spark.sparkContext.hadoopConfiguration)

// 递归遍历所有文件并累加大小
val totalSize = spark.sparkContext.parallelize(Seq(new Path(deltaPath)))
  .flatMap(path => {
    val statuses = fs.listFiles(path, true)
    Iterator.continually(if (statuses.hasNext) statuses.next() else null).takeWhile(_ != null)
  })
  .filter(!_.isDirectory)
  .map(_.getLen())
  .sum()

println(s"Delta表完整目录大小(含历史):${totalSize / 1024 / 1024} MB")

2. 仅当前快照的活跃数据大小(符合业务逻辑时使用)

如果只需要当前可用的数据大小,微软文档的方法1是准确的,但需确认业务需求是否排除历史文件:

import com.databricks.sql.transaction.tahoe._
val deltaLog = DeltaLog.forTable(spark, "abfss://<container>@<storage-account>.dfs.core.windows.net/<delta-table-path>")
println(s"当前快照活跃数据大小:${deltaLog.snapshot.sizeInBytes / 1024 / 1024} MB")

二、普通文件夹(非Delta表)大小统计

同样采用Spark分布式遍历,或借助Azure CLI调用ADLS Gen2 API实现高效统计:

方法1:Spark分布式统计

import org.apache.hadoop.fs._
val folderPath = "abfss://<container>@<storage-account>.dfs.core.windows.net/<folder-path>"
val fs = FileSystem.get(spark.sparkContext.hadoopConfiguration)

val totalSize = spark.sparkContext.parallelize(Seq(new Path(folderPath)))
  .flatMap(path => {
    val statuses = fs.listFiles(path, true)
    Iterator.continually(if (statuses.hasNext) statuses.next() else null).takeWhile(_ != null)
  })
  .filter(!_.isDirectory)
  .map(_.getLen())
  .sum()

println(s"文件夹总大小:${totalSize / 1024 / 1024 / 1024} GB")

方法2:Azure CLI快速统计(无需Spark集群)

通过ADLS Gen2的API并行获取文件元数据,适合轻量场景:

az storage blob list \
  --account-name <storage-account> \
  --container-name <container> \
  --prefix <folder-path> \
  --query "[].properties.contentLength" \
  --output tsv | awk '{sum+=$1} END {print sum / 1024 / 1024 / 1024 " GB"}'

方案对比

方案适用场景优势劣势
Spark分布式遍历超大目录(10亿级文件)分布式并行,速度快,支持任意存储路径需要Spark集群资源
Azure CLI命令中小目录、无Spark集群时使用轻量,无需集群,API调用效率高超大目录下需分页处理
DeltaLog快照统计仅需当前活跃数据的Delta表直接读取元数据,最快且准确不包含历史文件

内容的提问来源于stack exchange,提问作者Joshua Stafford

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 16:09:15