使用pytest-recording测试S3下载时触发IncompleteReadError问题
关于pytest-recording测试S3文件读取时的IncompleteReadError问题
问题场景
我在使用pytest-recording测试S3连接与数据下载功能时遇到以下问题:
- 仅执行下载测试(无上传操作),使用生产环境变量
- REPL环境中,相同代码可正常连接S3实例、读取并下载数据
- 运行
pytest --record-mode=once时,只要保留test_read_files_pandas测试用例,无论是否删除已有cassettes,都会触发botocore.exceptions.IncompleteReadError;注释该用例后所有测试均可正常通过
待测试原函数
def s3_connect_get_files( validated_target_params: dict, ) -> Tuple[s3fs.S3FileSystem, List[str]]: """ connects to our s3 instance, returning our bucket s3fs object and a list of the files in the target infile directory specified in our data model for the target. returns: - Tuple(the_bucket (s3fs obj), files (list) ) raises: - FileNotFoundError if we can't find the dir on our s3 """ try: the_bucket = s3fs.S3FileSystem( key=AWS_ACCESS_KEY_ID, secret=AWS_SECRET_ACCESS_KEY, client_kwargs={"endpoint_url": validated_target_params["endpoint_url"]}, ) files = the_bucket.glob( f"{validated_target_params['in_path']}/" f"{validated_target_params['glob_pattern']}" ) return the_bucket, files except FileNotFoundError as e: logger.error("could not connect to s3. check credentials!") logger.error(f"original error type: {type(e).__name__}") logger.error(f"original error message: {e}") raise def read_files( the_bucket: s3fs.S3FileSystem, files: List[str], validated_target_params: dict ) -> dict: """ reads all files into memory raises: - NotImplementedError; if we encounter a reader type we haven't defined yet. """ records = {} logger.info("Now reading data to be validated and de-duped.") for file in files: if ( validated_target_params["reader"].value == "pandas" ): # we need to call value as we're using an Enum try: df = pd.read_csv(the_bucket.open(file)) file_last_modified = the_bucket.info(file).get("LastModified") df["file_last_modified"] = file_last_modified records[file] = { "data": df, "last_modified_at": file_last_modified, } logger.info(f"Loaded {file} with {len(df)} rows using pandas") except Exception as e: logger.error(f"original error type: {type(e).__name__}") logger.error(f"original error message: {e}") raise else: raise NotImplementedError( f"{validated_target_params['reader']} is not yet implemented as a reader." ) return records
测试代码
@pytest.mark.vcr() def test_s3_connect_get_files(validated_target_params) -> None: """ test s3 connection and file retrieval """ the_bucket, files = s3_connect_get_files(validated_target_params) assert isinstance(the_bucket, S3FileSystem) assert isinstance(files, list) @pytest.mark.vcr() def test_read_files_pandas(validated_target_params) -> None: """ test reading files using pandas """ the_bucket, files = s3_connect_get_files(validated_target_params) records = read_files(the_bucket, files, validated_target_params) assert isinstance(records, dict) assert len(records) == len(files) assert all(isinstance(df["data"], pd.DataFrame) for df in records.values())
错误信息
E botocore.exceptions.IncompleteReadError: 0 read, but total bytes expected is 6163243. .venv/lib/python3.11/site-packages/aiobotocore/response.py:125: IncompleteReadError
解决建议
1. 调整S3文件读取方式,避免流式录制问题
the_bucket.open(file)返回的流式对象可能导致vcr无法完整录制大文件的响应数据。修改read_files中的读取逻辑,一次性读取完整内容后再传给pandas:
# 导入io模块 import io # 替换原df = pd.read_csv(the_bucket.open(file)) with the_bucket.open(file, 'rb') as f: content = f.read() df = pd.read_csv(io.BytesIO(content))
一次性读取完整内容后,vcr可以完整录制请求和响应,避免IncompleteRead异常。
2. 优化vcr录制配置
在@pytest.mark.vcr()中添加配置,忽略可能导致校验失败的请求头,或调整录制模式:
@pytest.mark.vcr(ignore_headers=['Content-Length'], record_mode='new_episodes')
也可以设置更精确的请求匹配规则:
@pytest.mark.vcr(match_on=['method', 'uri', 'body'])
3. 复用S3连接fixture,避免重复请求
当前两个测试用例重复调用s3_connect_get_files,可能导致录制的cassette冲突。用fixture复用S3连接和文件列表:
@pytest.fixture @pytest.mark.vcr() def s3_bucket_and_files(validated_target_params): return s3_connect_get_files(validated_target_params) def test_s3_connect_get_files(s3_bucket_and_files): the_bucket, files = s3_bucket_and_files assert isinstance(the_bucket, S3FileSystem) assert isinstance(files, list) @pytest.mark.vcr() def test_read_files_pandas(s3_bucket_and_files, validated_target_params): the_bucket, files = s3_bucket_and_files records = read_files(the_bucket, files, validated_target_params) assert isinstance(records, dict) assert len(records) == len(files) assert all(isinstance(df["data"], pd.DataFrame) for df in records.values())
4. 排查文件大小与服务稳定性
尝试用小文件测试,确认是否是大文件录制导致的问题;如果是自托管S3服务,检查服务端是否存在响应截断的情况。
内容的提问来源于stack exchange,提问作者nikUoM
相关产品推荐
相关产品推荐

