如何用Pandas读取含类列表路径结构的Azure Blob文件?
解决Pandas读取ADLS文件时路径特殊字符引发的IndexError问题
当使用Pandas的read_csv读取Azure Data Lake Storage(ADLS)中的文件时,若路径包含Type - ['D']这类带方括号的特殊字符,会被底层的fsspec工具误解析为列表匹配语法,最终触发IndexError: list index out of range错误。
错误原因
fsspec会将路径中的方括号[]识别为通配符或列表匹配规则,尝试将路径拆分为多个文件的列表,但实际不存在对应匹配项,导致索引越界。
解决办法
1. URL编码特殊字符
将路径中的方括号替换为URL编码:[替换为%5B,]替换为%5D,避免fsspec解析为语法符号。修改后的代码:
import pandas as pd datalake_connection_string = "<connection_string_for_the_container>" data = pd.read_csv( "abfs://container_name@storage_account_name.blob.core.windows.net/OutputFiles/CodeOutputs/Type - %5B'D'%5D/SOLUTIONS/summary.csv", storage_options={"connection_string": datalake_connection_string} )
2. 使用fsspec转义语法
fsspec支持用双括号表示字面量的方括号:[[对应[,]]对应]。修改路径中的特殊部分:
import pandas as pd datalake_connection_string = "<connection_string_for_the_container>" data = pd.read_csv( "abfs://container_name@storage_account_name.blob.core.windows.net/OutputFiles/CodeOutputs/Type - [[']D[']]/SOLUTIONS/summary.csv", storage_options={"connection_string": datalake_connection_string} )
3. 直接通过ADLS文件系统对象读取
先创建AzureBlobFileSystem对象,直接通过文件系统打开文件,绕开路径解析问题,这种方法更可靠:
import pandas as pd from adlfs import AzureBlobFileSystem datalake_connection_string = "<connection_string_for_the_container>" # 初始化ADLS文件系统 fs = AzureBlobFileSystem(connection_string=datalake_connection_string) # 打开文件并读取CSV with fs.open("container_name/OutputFiles/CodeOutputs/Type - ['D']/SOLUTIONS/summary.csv") as file: data = pd.read_csv(file)
内容的提问来源于stack exchange,提问作者Saipraneeth
相关产品推荐
相关产品推荐

