You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Boost Accumulators统计大于分位数的元素数量

解决方案

首先明确:Boost Accumulators没有直接提供统计大于指定值元素数量的内置提取器,但可以通过两种方式实现需求,前提是你的accumulator保留了所有原始样本(如果使用的是p_square_quantile这类仅计算分位数、不存储样本的策略,无法事后统计,因为没有原始数据)。

方法1:存储所有样本并统计

在定义accumulator_set时,添加ordered_sample统计项来存储所有元素,之后直接遍历样本统计大于分位数的数量。

示例代码

#include <boost/accumulators/accumulators.hpp>
#include <boost/accumulators/statistics.hpp>
#include <boost/accumulators/statistics/ordered_sample.hpp>
#include <algorithm>
#include <iostream>

namespace ba = boost::accumulators;

// 定义包含count、quantile、ordered_sample的accumulator
using Accumulator = ba::accumulator_set<double,
    ba::features<
        ba::tag::count,
        ba::tag::quantile,
        ba::tag::ordered_sample
    >
>;

int main() {
    // 初始化accumulator,指定要计算的分位数概率
    Accumulator acc(ba::quantile_probabilities = std::vector<double>{0.99});

    // 模拟添加数据
    for (int i = 0; i < 1000; ++i) {
        acc(static_cast<double>(i));
    }

    // 获取0.99分位数q
    double q = ba::quantile(acc, ba::quantile_probability = 0.99);

    // 提取所有样本并统计大于q的数量
    const auto& samples = ba::ordered_sample(acc);
    auto count_gt_q = std::count_if(samples.begin(), samples.end(),
        [q](double val) { return val > q; });

    std::cout << "总元素数: " << ba::count(acc) << "\n";
    std::cout << "0.99分位数: " << q << "\n";
    std::cout << "大于q的元素数: " << count_gt_q << "\n";

    return 0;
}

方法2:自定义带参数的提取器

如果希望用更贴合Boost Accumulators风格的方式实现,可以自定义一个接受阈值参数的提取器,本质也是基于存储的样本进行统计。

示例代码

#include <boost/accumulators/accumulators.hpp>
#include <boost/accumulators/statistics.hpp>
#include <boost/accumulators/statistics/ordered_sample.hpp>
#include <boost/accumulators/framework/extractor.hpp>
#include <boost/range/algorithm/count_if.hpp>
#include <iostream>

namespace ba = boost::accumulators;

// 定义自定义标签
namespace my_tags {
    struct count_greater_than {};
}

// 实现提取器逻辑
namespace boost::accumulators {
    template<typename AccumulatorSet>
    std::size_t extract(my_tags::count_greater_than, const AccumulatorSet& acc, double threshold) {
        const auto& samples = extract(tag::ordered_sample{}, acc);
        return boost::count_if(samples, [threshold](double val) { return val > threshold; });
    }
}

int main() {
    Accumulator acc(ba::quantile_probabilities = std::vector<double>{0.99});

    // 模拟添加数据
    for (int i = 0; i < 1000; ++i) {
        acc(static_cast<double>(i));
    }

    double q = ba::quantile(acc, ba::quantile_probability = 0.99);
    // 使用自定义提取器统计大于q的元素数
    std::size_t count_gt_q = ba::extract(my_tags::count_greater_than{}, acc, q);

    std::cout << "大于q的元素数: " << count_gt_q << "\n";
    return 0;
}

关键注意事项

  • 如果使用p_square_quantile、adaptive_quantile等不存储原始样本的分位数策略,无法事后统计大于q的元素数量,因为accumulator仅保留计算分位数的必要状态,没有原始数据。这种场景下只能在累加元素时实时维护计数器,但无法适配动态变化的q值。
  • ordered_sample会存储所有样本,适合小数据量场景;如果数据量极大,建议在累加时提前跟踪目标分位数对应的计数,或改用其他统计库(如Boost Histogram)。

内容的提问来源于stack exchange,提问作者SNJ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 06:15:08