如何使用Python hypothesis库生成列值存在依赖的Pandas DataFrame
基于hypothesis生成带列依赖的Pandas DataFrame解决方案
问题原因
你当前写法中使用的st.shared(id1)的作用是让id2策略全局共享id1生成的单个值,而非每行维度共享对应行的id1值,因此会出现所有行的id2均为同一固定值的问题,不符合列联动的需求。
解决方案
方案1:行级策略定义(适合复杂行内多列依赖)
通过data_frames的rows参数定义每行的字段依赖规则,可灵活实现多列的联动逻辑:
from hypothesis import strategies as st from hypothesis.extra.pandas import data_frames, column, range_indexes def create_dataframe(): # 定义单行列值联动规则 row_strategy = st.fixed_dictionaries({ "id1": st.integers(), }).map(lambda d: (d["id1"], d["id1"] * 2)) return data_frames( index=range_indexes(min_size=10, max_size=100), columns=[ column(name="id1", dtype=int), column(name="id2", dtype=int) ], rows=row_strategy ).filter(lambda df: df["id1"].is_unique)
方案2:基础DF扩展(适合整列计算类依赖,更推荐)
如果列依赖是整列的简单计算逻辑,可先生成基础字段的DataFrame,再通过map方法新增依赖列,性能更优且逻辑更直观,不需要额外去重过滤:
from hypothesis import strategies as st from hypothesis.extra.pandas import data_frames, column, range_indexes def create_dataframe(): # 先生成含唯一id1的基础DataFrame base_df = data_frames( index=range_indexes(min_size=10, max_size=100), columns=[column(name="id1", elements=st.integers(), unique=True, dtype=int)] ) # 新增联动列id2 return base_df.map(lambda df: df.assign(id2=df["id1"] * 2))
两种方案生成的DataFrame都能保证每行的id2等于对应行id1的两倍,满足列联动需求。
内容的提问来源于stack exchange,提问作者Jeff Harrison
相关产品推荐
相关产品推荐

