如何用Hypothesis的one_of为Pandas列生成多种数据类型?
解决方法
问题核心在于hs_pd.column的dtype参数无法直接接收hs.one_of返回的策略对象,需要换一种方式生成多类型列:
方案1:为每个列组合不同dtype的列策略
直接为每个列创建包含所有目标dtype的列策略集合,用hs.one_of包裹不同dtype对应的hs_pd.column:
from hypothesis import given import hypothesis.strategies as hs import hypothesis.extra.numpy as hs_np import hypothesis.extra.pandas as hs_pd import numpy as np import pandas as pd import pandera as pda import pytest data_schema = pda.DataFrameSchema(...) def non_float64_column(name: str) -> hs.SearchStrategy[pd.Series]: # 为指定列名生成所有非float64类型的列策略 return hs.one_of( hs_pd.column(name, dtype=hs_np.integer_dtypes()), hs_pd.column(name, dtype=hs_np.complex_number_dtypes()), hs_pd.column(name, dtype=hs_np.datetime64_dtypes()), hs_pd.column(name, dtype=hs_np.timedelta64_dtypes()), ) @given( hs_pd.data_frames([ non_float64_column("x"), non_float64_column("y"), non_float64_column("z"), ]) ) def test_invalid(df: pd.DataFrame) -> None: r"""Test that the schema does not pass invalid data.""" # 注意要指定完整的SchemaError路径 with pytest.raises(pda.errors.SchemaError): _ = data_schema(df)
方案2:用hs.builds动态生成列策略
通过hs.builds结合dtype策略,动态生成对应类型的列:
def non_float64_column(name: str) -> hs.SearchStrategy[pd.Series]: dtype_strategy = hs.one_of( hs_np.integer_dtypes(), hs_np.complex_number_dtypes(), hs_np.datetime64_dtypes(), hs_np.timedelta64_dtypes(), ) # 根据随机生成的dtype创建对应的列 return hs.builds( lambda dtype: hs_pd.column(name, dtype=dtype).example(), dtype=dtype_strategy )
关键说明
hs_pd.column的dtype参数仅支持直接传入具体dtype或能生成单一dtype的策略,无法直接解析hs.one_of的多类型策略集合。- 方案1更直观,每个列会随机从指定的dtype集合中选择一种类型生成数据,完全满足测试需求。
内容的提问来源于stack exchange,提问作者shadowtalker
相关产品推荐
相关产品推荐

