You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Azure Databricks中递归计算Azure Data Lake的总存储大小?

在Azure Databricks中递归计算ADLS存储大小的方法

以下是两种实用的方法,可在Azure Databricks中递归统计ADLS各文件夹及子文件夹的存储大小:

方法一:使用dbutils.fs递归遍历(适合小到中等规模数据)

通过自定义递归函数遍历所有文件和文件夹,累加文件大小并按层级统计:

def calculate_folder_size(path: str, size_dict: dict = None) -> dict:
    if size_dict is None:
        size_dict = {}
    
    # 获取当前路径下的所有文件和子文件夹
    items = dbutils.fs.ls(path)
    
    total_size = 0
    for item in items:
        if item.isFile():
            # 累加文件大小(单位:字节)
            total_size += item.size
        else:
            # 递归处理子文件夹
            sub_size = calculate_folder_size(item.path, size_dict)
            total_size += sub_size
    
    # 记录当前文件夹的总大小,转换为更易读的GB单位
    size_dict[path] = round(total_size / (1024**3), 2)
    return total_size

# 替换为你的ADLS路径(示例:abfss://container@storageaccount.dfs.core.windows.net/root/)
root_path = "abfss://your-container@your-storage-account.dfs.core.windows.net/your-root-folder/"
folder_sizes = calculate_folder_size(root_path)

# 打印结果
print("各文件夹存储大小(单位:GB):")
for folder, size in folder_sizes.items():
    print(f"{folder}: {size} GB")

方法二:使用Spark文件统计API(适合大规模数据)

利用Spark的底层Hadoop文件系统API批量获取文件元数据,并行处理效率更高:

from pyspark.sql import functions as F

# 替换为你的ADLS根路径
root_path = "abfss://your-container@your-storage-account.dfs.core.windows.net/your-root-folder/"

# 递归获取所有文件的元数据
files_df = spark.read.format("binaryFile")\
    .option("recursiveFileLookup", "true")\
    .load(root_path)

# 提取文件夹路径,按文件夹分组计算总大小
folder_size_df = files_df.withColumn("folder_path", F.regexp_extract("path", r"(.*)/[^/]+$", 1))\
    .groupBy("folder_path")\
    .agg(F.sum("length").alias("total_size_bytes"))\
    .withColumn("total_size_gb", round(F.col("total_size_bytes") / (1024**3), 2))\
    .orderBy("total_size_gb", ascending=False)

# 展示结果(可在Databricks中直接查看可视化表格)
display(folder_size_df)

注意事项

  • 确保你的Databricks集群已配置好ADLS访问权限(支持服务主体、托管标识或SAS令牌三种方式)
  • ABFS路径格式需正确:abfss://<container-name>@<storage-account-name>.dfs.core.windows.net/<folder-path>
  • 大规模数据场景优先选方法二,Spark并行处理能显著提升统计效率

内容的提问来源于stack exchange,提问作者Shikha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 16:52:33