如何列出不同订阅下Azure Data Lake Gen2存储账户内的所有文件与子目录
问题根因
- Python标准库的
os.walk()仅支持遍历本地文件系统,无法识别abfss这类ADLS Gen2的专有存储协议,无法读取远程Azure存储上的目录和文件,因此返回空列表。
适配abfss的正确实现(Spark环境)
前提:你已经完成对应Azure订阅存储账户的Spark访问权限配置(如Service Principal、SAS令牌等配置已写入Spark上下文),可正常读写abfss路径下的文件。
方案1:基于Hadoop FileSystem API实现(通用适配所有Spark环境)
from fnmatch import fnmatch from pyspark.sql import SparkSession from org.apache.hadoop.fs import Path as HadoopPath # 初始化SparkSession(如果当前环境已初始化可跳过) spark = SparkSession.builder.getOrCreate() hadoop_conf = spark._jsc.hadoopConfiguration() root_path = "abfss://xxxxcontainerxxxx@xxxxxstorageaccountxxxx.dfs.core.windows.net/staging/" files_list = [] pattern = "*.*" # 构造Hadoop路径对象 root_hadoop_path = HadoopPath(root_path) fs = root_hadoop_path.getFileSystem(hadoop_conf) # 递归遍历所有文件(第二个参数True代表递归遍历子目录) file_statuses = fs.listFiles(root_hadoop_path, True) while file_statuses.hasNext(): file_status = file_statuses.next() # 跳过目录,只保留文件 if not file_status.isDirectory(): file_full_path = file_status.getPath().toString() file_name = file_status.getPath().getName() # 按规则过滤文件名 if fnmatch(file_name.upper(), pattern.upper()): files_list.append(file_full_path) print(files_list)
方案2:Databricks环境简化实现(仅适用Databricks)
如果使用Databricks运行代码,可直接调用dbutils工具实现遍历:
from fnmatch import fnmatch root_path = "abfss://xxxxcontainerxxxx@xxxxxstorageaccountxxxx.dfs.core.windows.net/staging/" files_list = [] pattern = "*.*" # 递归遍历目录 def list_files_recursive(path): for item in dbutils.fs.ls(path): if item.isDir(): list_files_recursive(item.path) else: file_name = item.name.split('/')[-1] if fnmatch(file_name.upper(), pattern.upper()): files_list.append(item.path) list_files_recursive(root_path) print(files_list)
注意事项
- 若执行后仍无结果,首先检查Spark权限配置是否生效:可先尝试用
spark.read.text(root_path + "测试文件路径")读取单个已知存在的文件,验证读写权限是否正常。 - 存储账户防火墙、私有端点配置也可能导致访问失败,需确认运行Spark的集群网络可正常访问目标存储账户。
内容的提问来源于stack exchange,提问作者nYuker_98 D
相关产品推荐
相关产品推荐

