Hypothesis测试框架:概率是否应在框架外管理?及列表生成问题
需求背景
需要生成一个列表,元素及对应权重如下:
- "a": 10
- "b": 20
- "c": 25
- "d": 35
列表长度由伯努利试验决定:每次试验失败概率为0.1(对应期望列表长度为9),一旦试验失败就终止列表生成。
初始依赖代码:
from hypothesis import strategies as st import random import time import math
初始实现及性能问题
以下是初始实现的Hypothesis策略:
@st.composite def choices_bernoulli(draw, population, weights, failure_weight): """ 模拟首次失败即终止的伯努利过程,失败概率由failure_weight指定。 """ random = draw(st.randoms()) fail = object() results = [] while True: choice = random.choices( (*population, fail), (*weights, failure_weight), k=1 )[0] if choice is fail: break results.append(choice) return results choices_bernoulli("abcd", [10,20,25,35], 10)
该策略生成100个样本耗时约35秒,性能很差。同时存在两个随机过程:一是从元素池加权采样元素,二是通过伯努利试验生成列表长度,目前纠结这两个过程是否应该放在Hypothesis框架外部实现。
关于Hypothesis使用的思考
Hypothesis的核心能力是在测试数据的搜索空间中采样,理想状态下能覆盖所有可能的组合,因此似乎不需要手动指定选择概率;但搜索空间通常过大,此时有两种方案可选:
- 均匀采样:依赖Hypothesis的收缩机制、
one_of优先小空间等特性,目标是覆盖所有代码路径 - 指定概率引导采样:引导Hypothesis探索更常见的输入,但多数bug源于意外输入而非常见输入的细微变化,因此指定概率并非良策;此外Hypothesis会操纵随机生成器,导致手动指定的概率失效(可通过
use_true_random参数规避)
对比random.choices(真假随机模式)与Hypothesis的st.sampled_from生成的列表长度分布,发现Hypothesis不会对random.choices的输出进行收缩,且初始策略生成的列表平均长度仅3-4,远低于期望的9:
random.choices: false random random.choices: true random strategies.sampled_from 0: *********************** 0: 0: ************************** 1: ******************* 1: 1: ********************** 2: ***************************** 2: ** 2: ********************** 3: ************************ 3: **** 3: ************************** 4: ******************** 4: ***** 4: *************************** 5: ************************** 5: ******** 5: ************************ 6: ********************* 6: ************* 6: ************************** 7: ********************** 7: ***************** 7: *************************** 8: ************************* 8: ************************** 8: ***************************** 9: ********************** 9: ***************************** 9: *********************
放弃概率后的尝试方案
既然决定放弃手动指定概率,转而采用均匀采样搜索空间,但对于仅由概率定义的列表长度,尝试了两种方案:
方案1:基于均值的均匀长度采样
尝试丢弃二项分布,生成符合原统计特征(均值)的数据:从与原负二项分布均值相同的均匀分布中采样列表长度,但该策略生成的列表平均长度仅4-5,未达到预期:
@st.composite def choices_bernoulli_mean(draw, population, failure_weight): fail_norm = failure_weight / 100 expected_k = 1 * (1 - fail_norm) / fail_norm max_size = int(expected_k * 2) return draw(st.lists( st.sampled_from(population), min_size=0, max_size=max_size ))
方案2:直接从负二项分布采样长度
尝试直接从负二项分布采样得到列表长度,再生成固定长度的列表,但遇到random.uniform分布不均匀的问题(替换为st.floats策略也存在同样问题),容易生成极端长列表,导致生成缓慢甚至触发Unsatisfiable超时:
@st.composite def choices_bernoulli_dist(draw, population, failure_weight): def inverse(i, p): if i <= 0 or i > p: raise ValueError return math.floor(math.log(i / p) / math.log(1 - p)) p = failure_weight / 100 random = draw(st.randoms(use_true_random=False)) i = random.uniform(0, p) hypothesis.assume(0 < i <= p) # 另一种取值方式: #i = draw(st.floats( # min_value=0, max_value=p, exclude_min=True, exclude_max=False #)) k = inverse(i, p) return draw(st.lists( st.sampled_from(population), min_size=k, max_size=k ))
当前困境与求助
目前面临两种选择:
- 一是使用真随机数生成器采样负二项分布
- 二是接受Hypothesis的随机生成器,自定义列表长度的上下限,但自定义限界又回到了元素采样时的外部分布问题;且即使不限界,随机过程也会干扰Hypothesis的搜索空间采样
同时均匀分布不可行(99分位数的列表长度为43,过长),因此考虑只覆盖75%(对应长度13)或50%(对应长度6)的样本,但这又引入了新的限界。或许需要Hypothesis驱动的分布(先尝试少量大值再收缩),但不确定该如何选择,寻求具体思路与解决方案。
此外,性能瓶颈主要在于st.sampled_from或random.choices的调用,列表生成本身的开销极低。
内容的提问来源于stack exchange,提问作者user19087

