You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas读取ADLS中Parquet文件遇错误,求解决方案

解决Pandas读取ADLS Gen2 Parquet文件的协议与认证问题

是否必须通过PySpark读取后转Pandas DataFrame?

不需要。Pandas配合adlfs库就能直接读取ADLS Gen2上的Parquet文件,之前的ValueError: Protocol not known: abfss是因为Pandas默认不支持abfss协议,adlfs作为文件系统扩展可补上该支持。

解决ClientAuthenticationError认证失败问题

出现该错误是因为adlfs缺少访问ADLS Gen2的合法认证信息,以下是几种常用解决方式:

方式1:代码内传入认证参数

在pd.read_parquet中通过storage_options指定存储账户信息:

import pandas as pd

parquet_path = 'abfss://<容器名>@<存储账户名>.dfs.core.windows.net/abc.parquet'
df = pd.read_parquet(
    parquet_path,
    engine='pyarrow',
    storage_options={
        'account_name': '你的存储账户名',
        'account_key': '你的存储账户密钥'
    }
)

方式2:设置环境变量(推荐,避免硬编码敏感信息)

在终端或系统环境中配置变量:

# Linux/macOS
export AZURE_STORAGE_ACCOUNT="你的存储账户名"
export AZURE_STORAGE_KEY="你的存储账户密钥"

# Windows(PowerShell)
$env:AZURE_STORAGE_ACCOUNT="你的存储账户名"
$env:AZURE_STORAGE_KEY="你的存储账户密钥"

之后代码无需额外参数即可读取:

import pandas as pd

parquet_path = 'abfss://<容器名>@<存储账户名>.dfs.core.windows.net/abc.parquet'
df = pd.read_parquet(parquet_path, engine='pyarrow')

方式3:Azure AD身份认证(服务主体/用户身份场景)

若使用Azure AD认证,可在storage_options中传入对应参数:

df = pd.read_parquet(
    parquet_path,
    engine='pyarrow',
    storage_options={
        'account_name': '你的存储账户名',
        'client_id': '服务主体ID',
        'client_secret': '服务主体密钥',
        'tenant_id': '租户ID'
    }
)

内容的提问来源于stack exchange,提问作者user19930511

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 21:39:17