You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pytest-recording测试S3下载时触发IncompleteReadError问题

关于pytest-recording测试S3文件读取时的IncompleteReadError问题

问题场景

我在使用pytest-recording测试S3连接与数据下载功能时遇到以下问题:

  • 仅执行下载测试(无上传操作),使用生产环境变量
  • REPL环境中,相同代码可正常连接S3实例、读取并下载数据
  • 运行pytest --record-mode=once时,只要保留test_read_files_pandas测试用例,无论是否删除已有cassettes,都会触发botocore.exceptions.IncompleteReadError;注释该用例后所有测试均可正常通过

待测试原函数

def s3_connect_get_files(
    validated_target_params: dict,
) -> Tuple[s3fs.S3FileSystem, List[str]]:
    """
    connects to our s3 instance, returning
    our bucket s3fs object and a list of the files
    in the target infile directory specified in
    our data model for the target.

    returns:
        - Tuple(the_bucket (s3fs obj), files (list) )

    raises:
        - FileNotFoundError if we can't find the dir
        on our s3
    """
    try:
        the_bucket = s3fs.S3FileSystem(
            key=AWS_ACCESS_KEY_ID,
            secret=AWS_SECRET_ACCESS_KEY,
            client_kwargs={"endpoint_url": validated_target_params["endpoint_url"]},
        )
        files = the_bucket.glob(
            f"{validated_target_params['in_path']}/"
            f"{validated_target_params['glob_pattern']}"
        )
        return the_bucket, files
    except FileNotFoundError as e:
        logger.error("could not connect to s3. check credentials!")
        logger.error(f"original error type: {type(e).__name__}")
        logger.error(f"original error message: {e}")
        raise

def read_files(
    the_bucket: s3fs.S3FileSystem, files: List[str], validated_target_params: dict
) -> dict:
    """
    reads all files into memory

    raises:
        - NotImplementedError; if we encounter a reader
        type we haven't defined yet.
    """
    records = {}

    logger.info("Now reading data to be validated and de-duped.")
    for file in files:
        if (
            validated_target_params["reader"].value == "pandas"
        ):  # we need to call value as we're using an Enum
            try:
                df = pd.read_csv(the_bucket.open(file))
                file_last_modified = the_bucket.info(file).get("LastModified")

                df["file_last_modified"] = file_last_modified

                records[file] = {
                    "data": df,
                    "last_modified_at": file_last_modified,
                }
                logger.info(f"Loaded {file} with {len(df)} rows using pandas")
            except Exception as e:
                logger.error(f"original error type: {type(e).__name__}")
                logger.error(f"original error message: {e}")
                raise

        else:
            raise NotImplementedError(
                f"{validated_target_params['reader']} is not yet implemented as a reader."
            )

    return records

测试代码

@pytest.mark.vcr()
def test_s3_connect_get_files(validated_target_params) -> None:
    """
    test s3 connection and file retrieval
    """
    the_bucket, files = s3_connect_get_files(validated_target_params)
    assert isinstance(the_bucket, S3FileSystem)
    assert isinstance(files, list)


@pytest.mark.vcr()
def test_read_files_pandas(validated_target_params) -> None:
    """
    test reading files using pandas
    """
    the_bucket, files = s3_connect_get_files(validated_target_params)
    records = read_files(the_bucket, files, validated_target_params)
    assert isinstance(records, dict)
    assert len(records) == len(files)
    assert all(isinstance(df["data"], pd.DataFrame) for df in records.values())

错误信息

E           botocore.exceptions.IncompleteReadError: 0 read, but total bytes expected is 6163243.

.venv/lib/python3.11/site-packages/aiobotocore/response.py:125: IncompleteReadError

解决建议

1. 调整S3文件读取方式,避免流式录制问题

the_bucket.open(file)返回的流式对象可能导致vcr无法完整录制大文件的响应数据。修改read_files中的读取逻辑,一次性读取完整内容后再传给pandas:

# 导入io模块
import io

# 替换原df = pd.read_csv(the_bucket.open(file))
with the_bucket.open(file, 'rb') as f:
    content = f.read()
    df = pd.read_csv(io.BytesIO(content))

一次性读取完整内容后,vcr可以完整录制请求和响应,避免IncompleteRead异常。

2. 优化vcr录制配置

在@pytest.mark.vcr()中添加配置,忽略可能导致校验失败的请求头,或调整录制模式:

@pytest.mark.vcr(ignore_headers=['Content-Length'], record_mode='new_episodes')

也可以设置更精确的请求匹配规则:

@pytest.mark.vcr(match_on=['method', 'uri', 'body'])

3. 复用S3连接fixture,避免重复请求

当前两个测试用例重复调用s3_connect_get_files,可能导致录制的cassette冲突。用fixture复用S3连接和文件列表:

@pytest.fixture
@pytest.mark.vcr()
def s3_bucket_and_files(validated_target_params):
    return s3_connect_get_files(validated_target_params)

def test_s3_connect_get_files(s3_bucket_and_files):
    the_bucket, files = s3_bucket_and_files
    assert isinstance(the_bucket, S3FileSystem)
    assert isinstance(files, list)

@pytest.mark.vcr()
def test_read_files_pandas(s3_bucket_and_files, validated_target_params):
    the_bucket, files = s3_bucket_and_files
    records = read_files(the_bucket, files, validated_target_params)
    assert isinstance(records, dict)
    assert len(records) == len(files)
    assert all(isinstance(df["data"], pd.DataFrame) for df in records.values())

4. 排查文件大小与服务稳定性

尝试用小文件测试,确认是否是大文件录制导致的问题;如果是自托管S3服务,检查服务端是否存在响应截断的情况。

内容的提问来源于stack exchange,提问作者nikUoM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 01:49:52