代理环境下Pandas读取S3存储桶Parquet文件路径越界报错
问题场景
需要读取S3目录下多个Schema一致的Parquet文件,未启用代理环境下代码可正常运行,启用代理后触发路径校验类报错。
报错信息
Traceback (most recent call last): File "script.py", line 158, in <module> df = pq.read_table(source=bucket_path, filesystem=s3).to_pandas() File "pyarrow\parquet\__init__.py", line 2737, in read_table dataset = _ParquetDatasetV2( File "\pyarrow\parquet\__init__.py", line 2351, in __init__ self._dataset = ds.dataset(path_or_paths, filesystem=filesystem, File "pyarrow\dataset.py", line 694, in dataset return _filesystem_dataset(source, **kwargs) File "pyarrow\dataset.py", line 447, in _filesystem_dataset factory = FileSystemDatasetFactory(fs, paths_or_selector, format, options) File "pyarrow\_dataset.pyx", line 2031, in pyarrow._dataset.FileSystemDatasetFactory.__init__ File "pyarrow\error.pxi", line 144, in pyarrow.lib.pyarrow_internal_check_status File "pyarrow\error.pxi", line 100, in pyarrow.lib.check_status pyarrow.lib.ArrowInvalid: GetFileInfo() yielded path 's3://test/files/part-00000-ed788628-0a6d-4ce9-b604-dd4c6ec75b6d-c000.snappy.parquet', which is outside base dir 's3://test/files/'
原有问题代码
import pyarrow.parquet as pq import s3fs bucket_path = 's3://test/files/' os.environ['https_proxy'] = 'http://proxy.com:4200' # proxies = { # 'https': f''http://proxy.com:4200', # 'http': f'http://proxy.com:4200' # } # s3 = s3fs.S3FileSystem(anon=False, config_kwargs={'proxies': proxies}) s3 = s3fs.S3FileSystem(anon=False) df = pq.read_table(source=bucket_path, filesystem=s3).to_pandas()
根因说明
该报错是代理环境下s3fs与pyarrow的路径校验逻辑冲突导致:
- 直接通过系统环境变量配置代理时,s3fs底层botocore客户端未正确加载代理配置,列举文件返回的路径格式与传入的基础路径做字符串前缀匹配时出现判定偏差
- 传入的基础路径末尾带斜杠,代理转发S3请求时路径规范化处理逻辑与直连环境不一致,触发pyarrow的目录边界校验
- 原注释中的代理配置代码存在语法错误,https代理值前多写了一个单引号,即使取消注释也无法正常生效
解决方法
按优先级依次尝试以下方案:
- 修正路径格式与代理配置
去掉bucket_path末尾的斜杠,不要依赖系统环境变量传递代理,显式给s3fs传入正确格式的代理配置,参考代码:
import pyarrow.parquet as pq import s3fs # 去掉路径末尾的斜杠 bucket_path = 's3://test/files' proxies = { 'https': 'http://proxy.com:4200', 'http': 'http://proxy.com:4200' } # 显式传入代理配置初始化s3客户端 s3 = s3fs.S3FileSystem(anon=False, config_kwargs={'proxies': proxies}) df = pq.read_table(source=bucket_path, filesystem=s3).to_pandas()
- 手动列举文件绕过内部校验
如果方案1仍报错,可先通过s3fs手动列举目录下所有parquet文件,将文件列表直接传入read_table,跳过pyarrow内部的目录前缀校验逻辑,参考代码:
# 接上述s3初始化逻辑 # 列举目录下所有parquet文件 file_list = s3.glob(f"{bucket_path}/*.parquet") # 拼接为完整S3路径 full_file_paths = [f"s3://{path}" for path in file_list] df = pq.read_table(source=full_file_paths, filesystem=s3).to_pandas()
内容的提问来源于stack exchange,提问作者HouKaide
相关产品推荐
相关产品推荐

