如何利用已获取的Credential通过PySpark读取ADLS Gen2文件
解决方案:用已有的InteractiveBrowserCredential通过PySpark读取ADLS Gen2文件
你之前的尝试失败主要是因为路径格式不正确以及Spark缺少ADLS Gen2的认证配置。以下是基于你已获取的credential对象的完整实现步骤:
1. 配置Spark的ADLS认证参数
需要将InteractiveBrowserCredential的认证信息传递给Spark的JVM底层,通过设置Spark配置项完成:
# 获取针对Azure存储服务的访问令牌 token = credential.get_token("https://storage.azure.com/.default") # 配置Spark连接ADLS Gen2的认证参数 spark.conf.set(f"fs.azure.account.auth.type.{storage_account_name}.dfs.core.windows.net", "OAuth") spark.conf.set(f"fs.azure.account.oauth.provider.type.{storage_account_name}.dfs.core.windows.net", "org.apache.hadoop.fs.azurebfs.oauth2.ClientCredsTokenProvider") spark.conf.set(f"fs.azure.account.oauth2.client.id.{storage_account_name}.dfs.core.windows.net", credential._client_id) spark.conf.set(f"fs.azure.account.oauth2.client.secret.{storage_account_name}.dfs.core.windows.net", token.token) spark.conf.set(f"fs.azure.account.oauth2.client.endpoint.{storage_account_name}.dfs.core.windows.net", f"https://login.microsoftonline.com/{tenant_id}/oauth2/token")
2. 构建正确的文件路径并读取
使用完整的abfss协议路径(默认启用SSL加密,替代未加密的abfs):
# 构建正确的ADLS文件路径 file_path_full = f"abfss://{file_system}@{storage_account_name}.dfs.core.windows.net/{file_path}" # 读取文件(匹配你pandas读取时的CSV格式参数) df_spark = spark.read.format("csv") \ .option("compression", "gzip") \ .option("sep", "\t") \ .option("header", "false") \ .option("encoding", "latin1") \ .load(file_path_full) # 验证读取结果 df_spark.show(5)
错误原因说明
- 第一个报错
abfs://project-data/myfile.txt.gz has invalid authority:路径缺少存储账户标识,正确格式必须包含@存储账户名.dfs.core.windows.net后缀。 - 第二个报错
Failure to initialize configuration:Spark未配置ADLS的OAuth认证参数,无法完成身份验证流程。
内容的提问来源于stack exchange,提问作者gaut
相关产品推荐
相关产品推荐

