使用Pandas读取ADLS中Parquet文件遇错误,求解决方案
解决Pandas读取ADLS Gen2 Parquet文件的协议与认证问题
是否必须通过PySpark读取后转Pandas DataFrame?
不需要。Pandas配合adlfs库就能直接读取ADLS Gen2上的Parquet文件,之前的ValueError: Protocol not known: abfss是因为Pandas默认不支持abfss协议,adlfs作为文件系统扩展可补上该支持。
解决ClientAuthenticationError认证失败问题
出现该错误是因为adlfs缺少访问ADLS Gen2的合法认证信息,以下是几种常用解决方式:
方式1:代码内传入认证参数
在pd.read_parquet中通过storage_options指定存储账户信息:
import pandas as pd parquet_path = 'abfss://<容器名>@<存储账户名>.dfs.core.windows.net/abc.parquet' df = pd.read_parquet( parquet_path, engine='pyarrow', storage_options={ 'account_name': '你的存储账户名', 'account_key': '你的存储账户密钥' } )
方式2:设置环境变量(推荐,避免硬编码敏感信息)
在终端或系统环境中配置变量:
# Linux/macOS export AZURE_STORAGE_ACCOUNT="你的存储账户名" export AZURE_STORAGE_KEY="你的存储账户密钥" # Windows(PowerShell) $env:AZURE_STORAGE_ACCOUNT="你的存储账户名" $env:AZURE_STORAGE_KEY="你的存储账户密钥"
之后代码无需额外参数即可读取:
import pandas as pd parquet_path = 'abfss://<容器名>@<存储账户名>.dfs.core.windows.net/abc.parquet' df = pd.read_parquet(parquet_path, engine='pyarrow')
方式3:Azure AD身份认证(服务主体/用户身份场景)
若使用Azure AD认证,可在storage_options中传入对应参数:
df = pd.read_parquet( parquet_path, engine='pyarrow', storage_options={ 'account_name': '你的存储账户名', 'client_id': '服务主体ID', 'client_secret': '服务主体密钥', 'tenant_id': '租户ID' } )
内容的提问来源于stack exchange,提问作者user19930511
相关产品推荐
相关产品推荐

