Azure Databricks读取ADLS嵌套子目录parquet文件夹报错如何解决
Azure Databricks 读取ADLS嵌套目录Parquet文件解决方案
问题根因
直接传入根路径读取报错,通常是因为Spark默认不会递归遍历所有层级子目录查找parquet文件,或是根目录下存在非parquet格式的干扰文件,也可能是schema推断失败导致。
可执行方案
方法1:开启递归遍历直接读取
添加recursiveFileLookup参数强制Spark遍历所有子目录,适合文件格式统一、schema一致的场景:
# 替换路径为你的ADLS实际路径,已挂载的话也可以用DBFS路径 df = spark.read.format("parquet") \ .option("recursiveFileLookup", "true") \ .option("inferSchema", "true") \ .load("abfss://<你的容器名>@<存储账户名>.dfs.core.windows.net/base_folder/filename/")
方法2:手动指定schema读取
如果开启递归后仍报schema相关错误,手动定义schema即可解决推断失败问题:
from pyspark.sql.types import StructType, StructField, StringType, IntegerType, TimestampType # 根据你的Parquet实际字段结构修改schema定义 custom_schema = StructType([ StructField("业务id", IntegerType(), nullable=True), StructField("内容字段", StringType(), nullable=True), StructField("生成时间", TimestampType(), nullable=True) ]) df = spark.read.format("parquet") \ .option("recursiveFileLookup", "true") \ .schema(custom_schema) \ .load("abfss://<你的容器名>@<存储账户名>.dfs.core.windows.net/base_folder/filename/")
方法3:过滤干扰文件+保留分区字段
如果根目录下存在非parquet文件,或者需要把路径中的年/月/日作为DataFrame的字段保留,使用以下写法:
from pyspark.sql.functions import input_file_name, split, col df = spark.read.format("parquet") \ .option("recursiveFileLookup", "true") \ .option("pathGlobFilter", "*.parquet") \ # 仅读取.parquet后缀的文件 .schema(custom_schema) \ .load("abfss://<你的容器名>@<存储账户名>.dfs.core.windows.net/base_folder/filename/") # 从文件路径中提取年、月、日分区字段 df = df.withColumn("file_full_path", input_file_name()) \ .withColumn("年", split(col("file_full_path"), "/").getItem(-4)) \ .withColumn("月", split(col("file_full_path"), "/").getItem(-3)) \ .withColumn("日", split(col("file_full_path"), "/").getItem(-2)) \ .drop("file_full_path")
前置校验
操作前请先确认ADLS访问权限正常:你可以选择将ADLS容器挂载到Databricks DBFS,或是在Spark配置中提前配置存储账户的访问密钥、SAS令牌或服务主体权限,避免权限类报错。
内容的提问来源于stack exchange,提问作者Sharyu Aadhatrao
相关产品推荐
相关产品推荐

