如何用pytest模拟pandas read_csv测试GCS文件读取函数?
问题描述
我有一个Python函数,以BigQuery存储桶名称和文件路径为输入,执行以下操作:
- 检查存储桶是否存在
- 检查文件是否在存储桶中
- 将文件读取为DataFrame并返回该DataFrame
函数代码如下:
def function_to_test(client, bucket_name, delimiter, full_path=None, header=None): try: bucket = client.get_bucket(bucket_name) assert bucket is not None except (gcp_exceptions.GoogleCloudError, AssertionError) as ex: raise AirflowFailException(f"Failed to access bucket: {bucket_name}") from ex try: blob = bucket.get_blob(full_path) assert blob is not None if blob.size == 0: print(f'File is empty') return None df = pd.read_csv(f"gs://{bucket_name}/{full_path}", sep=delimiter, dtype='str', header=header) return df except(gcp_exceptions.GoogleCloudError, AssertionError) as ex: raise AirflowFailException(f"Failed to retrieve blob from bucket: {bucket_name}") from ex
我正在尝试为该函数编写pytest测试,目前已有如下测试代码:
def test_func(mocker, generic_df): bucket_name = 'test_bucket' full_path = 'test_path' mock_client = mocker.patch('google.cloud.storage.Client', autospec=True) mock_bucket = mock_client.get_bucket(bucket_name) actual_df = function_to_test(client=mock_client,bucket_name=bucket_name, delimiter=',', full_path=full_path, header=0)
目前对存储桶的模拟已满足函数中的存储桶校验要求,但我无法实现对read_csv功能的模拟,导致DataFrame创建失败。请问是否有方法可以模拟该函数,从而同时模拟DataFrame?
解决方案
你需要同时完善Blob对象的模拟和pandas.read_csv的拦截,具体步骤如下:
完善Blob对象模拟
测试中仅模拟Bucket还不够,需要给get_blob返回一个带有非零size的Blob对象,避免触发空文件分支:# 模拟Blob并设置非零大小 mock_blob = mocker.Mock() mock_blob.size = 100 mock_bucket.get_blob.return_value = mock_blob模拟pandas.read_csv方法
使用mocker.patch直接拦截pandas.read_csv的调用,让它返回你预先准备的测试用DataFrame:# 拦截read_csv,返回测试DataFrame mock_read_csv = mocker.patch('pandas.read_csv') mock_read_csv.return_value = generic_df完整测试代码
整合后的测试代码如下,还可以添加断言验证结果和调用逻辑:def test_func(mocker, generic_df): bucket_name = 'test_bucket' full_path = 'test_path' # 模拟GCS Client mock_client = mocker.patch('google.cloud.storage.Client', autospec=True) mock_bucket = mock_client.get_bucket.return_value # 模拟Blob对象,设置非零大小 mock_blob = mocker.Mock() mock_blob.size = 100 mock_bucket.get_blob.return_value = mock_blob # 模拟read_csv返回测试DataFrame mock_read_csv = mocker.patch('pandas.read_csv') mock_read_csv.return_value = generic_df # 调用待测试函数 actual_df = function_to_test( client=mock_client, bucket_name=bucket_name, delimiter=',', full_path=full_path, header=0 ) # 断言结果匹配 pd.testing.assert_frame_equal(actual_df, generic_df) # 验证read_csv的调用参数是否正确 mock_read_csv.assert_called_once_with( f"gs://{bucket_name}/{full_path}", sep=',', dtype='str', header=0 )扩展空文件测试场景(可选)
你还可以测试空文件的分支逻辑:def test_func_empty_file(mocker): bucket_name = 'test_bucket' full_path = 'test_path' mock_client = mocker.patch('google.cloud.storage.Client', autospec=True) mock_bucket = mock_client.get_bucket.return_value # 设置Blob大小为0 mock_blob = mocker.Mock() mock_blob.size = 0 mock_bucket.get_blob.return_value = mock_blob result = function_to_test( client=mock_client, bucket_name=bucket_name, delimiter=',', full_path=full_path, header=0 ) assert result is None
内容的提问来源于stack exchange,提问作者nimgwfc
相关产品推荐
相关产品推荐

