You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Hypothesis测试框架:概率是否应在框架外管理?及列表生成问题

问题:基于Hypothesis生成符合伯努利终止规则的加权抽样列表

需求背景

需要生成一个列表,元素及对应权重如下:

  • "a": 10
  • "b": 20
  • "c": 25
  • "d": 35

列表长度由伯努利试验决定:每次试验失败概率为0.1(对应期望列表长度为9),一旦试验失败就终止列表生成。

初始依赖代码:

from hypothesis import strategies as st
import random
import time
import math

初始实现及性能问题

以下是初始实现的Hypothesis策略:

@st.composite
def choices_bernoulli(draw, population, weights, failure_weight):
    """
    模拟首次失败即终止的伯努利过程,失败概率由failure_weight指定。
    """
    random = draw(st.randoms())
    fail = object()
    results = []
    while True:
        choice = random.choices(
            (*population, fail), (*weights, failure_weight), k=1
        )[0]
        if choice is fail:
            break
        results.append(choice)
    return results

choices_bernoulli("abcd", [10,20,25,35], 10)

该策略生成100个样本耗时约35秒,性能很差。同时存在两个随机过程:一是从元素池加权采样元素,二是通过伯努利试验生成列表长度,目前纠结这两个过程是否应该放在Hypothesis框架外部实现。

关于Hypothesis使用的思考

Hypothesis的核心能力是在测试数据的搜索空间中采样,理想状态下能覆盖所有可能的组合,因此似乎不需要手动指定选择概率;但搜索空间通常过大,此时有两种方案可选:

  • 均匀采样:依赖Hypothesis的收缩机制、one_of优先小空间等特性,目标是覆盖所有代码路径
  • 指定概率引导采样:引导Hypothesis探索更常见的输入,但多数bug源于意外输入而非常见输入的细微变化,因此指定概率并非良策;此外Hypothesis会操纵随机生成器,导致手动指定的概率失效(可通过use_true_random参数规避)

对比random.choices(真假随机模式)与Hypothesis的st.sampled_from生成的列表长度分布,发现Hypothesis不会对random.choices的输出进行收缩,且初始策略生成的列表平均长度仅3-4,远低于期望的9:

random.choices: false random            random.choices: true random             strategies.sampled_from
 0: ***********************              0:                                      0: **************************
 1: *******************                  1:                                      1: **********************
 2: *****************************        2: **                                   2: **********************
 3: ************************             3: ****                                 3: **************************
 4: ********************                 4: *****                                4: ***************************
 5: **************************           5: ********                             5: ************************
 6: *********************                6: *************                        6: **************************
 7: **********************               7: *****************                    7: ***************************
 8: *************************            8: **************************           8: *****************************
 9: **********************               9: *****************************        9: *********************

放弃概率后的尝试方案

既然决定放弃手动指定概率,转而采用均匀采样搜索空间,但对于仅由概率定义的列表长度,尝试了两种方案:

方案1:基于均值的均匀长度采样

尝试丢弃二项分布,生成符合原统计特征(均值)的数据:从与原负二项分布均值相同的均匀分布中采样列表长度,但该策略生成的列表平均长度仅4-5,未达到预期:

@st.composite
def choices_bernoulli_mean(draw, population, failure_weight):
    fail_norm = failure_weight / 100
    expected_k = 1 * (1 - fail_norm) / fail_norm
    max_size = int(expected_k * 2)
    return draw(st.lists(
        st.sampled_from(population), min_size=0, max_size=max_size
    ))

方案2:直接从负二项分布采样长度

尝试直接从负二项分布采样得到列表长度,再生成固定长度的列表,但遇到random.uniform分布不均匀的问题(替换为st.floats策略也存在同样问题),容易生成极端长列表,导致生成缓慢甚至触发Unsatisfiable超时:

@st.composite
def choices_bernoulli_dist(draw, population, failure_weight):
    def inverse(i, p):
        if i <= 0 or i > p:
            raise ValueError
        return math.floor(math.log(i / p) / math.log(1 - p))
    p = failure_weight / 100

    random = draw(st.randoms(use_true_random=False))
    i = random.uniform(0, p)
    hypothesis.assume(0 < i <= p)

    # 另一种取值方式:
    #i = draw(st.floats(
    #    min_value=0, max_value=p, exclude_min=True, exclude_max=False
    #))

    k = inverse(i, p)
    return draw(st.lists(
        st.sampled_from(population), min_size=k, max_size=k
    ))

当前困境与求助

目前面临两种选择:

  • 一是使用真随机数生成器采样负二项分布
  • 二是接受Hypothesis的随机生成器,自定义列表长度的上下限,但自定义限界又回到了元素采样时的外部分布问题;且即使不限界,随机过程也会干扰Hypothesis的搜索空间采样

同时均匀分布不可行(99分位数的列表长度为43,过长),因此考虑只覆盖75%(对应长度13)或50%(对应长度6)的样本,但这又引入了新的限界。或许需要Hypothesis驱动的分布(先尝试少量大值再收缩),但不确定该如何选择,寻求具体思路与解决方案。

此外,性能瓶颈主要在于st.sampled_from或random.choices的调用,列表生成本身的开销极低。

内容的提问来源于stack exchange,提问作者user19087

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 08:02:04