You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让moto兼容pandas的read_parquet与to_parquet函数?

问题:Moto与Pandas的S3 Parquet读写兼容问题

我在为一个使用pd.read_parquet()的函数编写单元测试时遇到了问题,测试代码如下:

from moto import mock_aws
import pandas as pd
import pytest
import datetime as dt
import boto3
from my_module import foo

@pytest.fixture
def mock_df():
    cols = [
        "timestamp",
        "value"
    ]
    values = [
        [dt.datetime(2024, 1, 1, 0), 2.57],
        [dt.datetime(2024, 1, 1, 1), 1.41],
        [dt.datetime(2024, 1, 1, 2), 2.06],
    ]
    df = pd.DataFrame(values, columns=cols)
    return df


@mock_aws
def test_download(mock_df):
    bucket_name = "test-input-bucket"
    s3 = boto3.resource("s3", region_name="us-east-1")
    s3.create_bucket(Bucket=bucket_name)
    key1 = "s3://test-input-bucket/path/to/data.parquet"
    mock_df.to_parquet(key1) # 代码在此处失败
    foo() # 使用pd.read_parquet()

运行时抛出以下错误:

OSError: When initiating multiple part upload for key 'path/to/data.parquet'
in bucket 'test-input-bucket': AWS Error INVALID_ACCESS_KEY_ID during 
CreateMultipartUpload operation: The AWS Access Key Id you provided does not exist in our records.

使用s3_bucket.put_object(Key=key1, Body=mock_df.to_parquet())的方式可以正常工作,但直接调用Pandas的to_parquet或read_parquet都会触发上述错误。由于无法替换Pandas的函数,需要找到让Moto与这些函数兼容的方法。使用版本:boto3 1.28.64、botocore 1.31.64、moto 5.0.3。


解决方案

1. 显式指定S3端点与存储选项

Pandas的Parquet读写函数默认使用独立的S3客户端,不会自动复用Moto Mock的配置。需要显式指定Moto的本地端点,并传递任意非空的密钥对(Moto接受任意合法格式的密钥)。修改测试代码如下:

@mock_aws
def test_download(mock_df):
    bucket_name = "test-input-bucket"
    s3 = boto3.resource("s3", region_name="us-east-1")
    s3.create_bucket(Bucket=bucket_name)
    key1 = "s3://test-input-bucket/path/to/data.parquet"
    
    # 添加storage_options参数,指定Moto端点和凭证
    storage_options = {
        "client_kwargs": {
            "endpoint_url": "http://localhost:5000",  # Moto默认端口为5000,自定义端口需对应调整
            "aws_access_key_id": "test",
            "aws_secret_access_key": "test"
        }
    }
    
    mock_df.to_parquet(key1, storage_options=storage_options)
    # 若foo函数内的read_parquet也需要兼容,需同样传递storage_options
    # 可通过参数传递或测试中注入配置实现

2. 确保Mock覆盖范围正确

确认@mock_aws装饰器完整覆盖测试函数的所有逻辑,或使用上下文管理器包裹关键代码块,避免S3客户端在Mock生效前初始化:

def test_download(mock_df):
    bucket_name = "test-input-bucket"
    with mock_aws():
        s3 = boto3.resource("s3", region_name="us-east-1")
        s3.create_bucket(Bucket=bucket_name)
        key1 = "s3://test-input-bucket/path/to/data.parquet"
        storage_options = {
            "client_kwargs": {
                "endpoint_url": "http://localhost:5000",
                "aws_access_key_id": "test",
                "aws_secret_access_key": "test"
            }
        }
        mock_df.to_parquet(key1, storage_options=storage_options)
        foo()

原理说明

Pandas内部的S3客户端不会自动读取Moto Mock注入的临时凭证,而是尝试使用系统默认的凭证链(环境变量、本地配置文件等),导致认证失败。通过显式指定Moto的本地端点和测试用密钥,让Pandas的请求直接指向Mock服务,从而绕过真实AWS的认证逻辑。


内容的提问来源于stack exchange,提问作者user430953

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 10:42:56