如何让moto兼容pandas的read_parquet与to_parquet函数?
问题:Moto与Pandas的S3 Parquet读写兼容问题
我在为一个使用pd.read_parquet()的函数编写单元测试时遇到了问题,测试代码如下:
from moto import mock_aws import pandas as pd import pytest import datetime as dt import boto3 from my_module import foo @pytest.fixture def mock_df(): cols = [ "timestamp", "value" ] values = [ [dt.datetime(2024, 1, 1, 0), 2.57], [dt.datetime(2024, 1, 1, 1), 1.41], [dt.datetime(2024, 1, 1, 2), 2.06], ] df = pd.DataFrame(values, columns=cols) return df @mock_aws def test_download(mock_df): bucket_name = "test-input-bucket" s3 = boto3.resource("s3", region_name="us-east-1") s3.create_bucket(Bucket=bucket_name) key1 = "s3://test-input-bucket/path/to/data.parquet" mock_df.to_parquet(key1) # 代码在此处失败 foo() # 使用pd.read_parquet()
运行时抛出以下错误:
OSError: When initiating multiple part upload for key 'path/to/data.parquet' in bucket 'test-input-bucket': AWS Error INVALID_ACCESS_KEY_ID during CreateMultipartUpload operation: The AWS Access Key Id you provided does not exist in our records.
使用s3_bucket.put_object(Key=key1, Body=mock_df.to_parquet())的方式可以正常工作,但直接调用Pandas的to_parquet或read_parquet都会触发上述错误。由于无法替换Pandas的函数,需要找到让Moto与这些函数兼容的方法。使用版本:boto3 1.28.64、botocore 1.31.64、moto 5.0.3。
解决方案
1. 显式指定S3端点与存储选项
Pandas的Parquet读写函数默认使用独立的S3客户端,不会自动复用Moto Mock的配置。需要显式指定Moto的本地端点,并传递任意非空的密钥对(Moto接受任意合法格式的密钥)。修改测试代码如下:
@mock_aws def test_download(mock_df): bucket_name = "test-input-bucket" s3 = boto3.resource("s3", region_name="us-east-1") s3.create_bucket(Bucket=bucket_name) key1 = "s3://test-input-bucket/path/to/data.parquet" # 添加storage_options参数,指定Moto端点和凭证 storage_options = { "client_kwargs": { "endpoint_url": "http://localhost:5000", # Moto默认端口为5000,自定义端口需对应调整 "aws_access_key_id": "test", "aws_secret_access_key": "test" } } mock_df.to_parquet(key1, storage_options=storage_options) # 若foo函数内的read_parquet也需要兼容,需同样传递storage_options # 可通过参数传递或测试中注入配置实现
2. 确保Mock覆盖范围正确
确认@mock_aws装饰器完整覆盖测试函数的所有逻辑,或使用上下文管理器包裹关键代码块,避免S3客户端在Mock生效前初始化:
def test_download(mock_df): bucket_name = "test-input-bucket" with mock_aws(): s3 = boto3.resource("s3", region_name="us-east-1") s3.create_bucket(Bucket=bucket_name) key1 = "s3://test-input-bucket/path/to/data.parquet" storage_options = { "client_kwargs": { "endpoint_url": "http://localhost:5000", "aws_access_key_id": "test", "aws_secret_access_key": "test" } } mock_df.to_parquet(key1, storage_options=storage_options) foo()
原理说明
Pandas内部的S3客户端不会自动读取Moto Mock注入的临时凭证,而是尝试使用系统默认的凭证链(环境变量、本地配置文件等),导致认证失败。通过显式指定Moto的本地端点和测试用密钥,让Pandas的请求直接指向Mock服务,从而绕过真实AWS的认证逻辑。
内容的提问来源于stack exchange,提问作者user430953
相关产品推荐
相关产品推荐

