使用Boost Accumulators统计大于分位数的元素数量
解决方案
首先明确:Boost Accumulators没有直接提供统计大于指定值元素数量的内置提取器,但可以通过两种方式实现需求,前提是你的accumulator保留了所有原始样本(如果使用的是p_square_quantile这类仅计算分位数、不存储样本的策略,无法事后统计,因为没有原始数据)。
方法1:存储所有样本并统计
在定义accumulator_set时,添加ordered_sample统计项来存储所有元素,之后直接遍历样本统计大于分位数的数量。
示例代码
#include <boost/accumulators/accumulators.hpp> #include <boost/accumulators/statistics.hpp> #include <boost/accumulators/statistics/ordered_sample.hpp> #include <algorithm> #include <iostream> namespace ba = boost::accumulators; // 定义包含count、quantile、ordered_sample的accumulator using Accumulator = ba::accumulator_set<double, ba::features< ba::tag::count, ba::tag::quantile, ba::tag::ordered_sample > >; int main() { // 初始化accumulator,指定要计算的分位数概率 Accumulator acc(ba::quantile_probabilities = std::vector<double>{0.99}); // 模拟添加数据 for (int i = 0; i < 1000; ++i) { acc(static_cast<double>(i)); } // 获取0.99分位数q double q = ba::quantile(acc, ba::quantile_probability = 0.99); // 提取所有样本并统计大于q的数量 const auto& samples = ba::ordered_sample(acc); auto count_gt_q = std::count_if(samples.begin(), samples.end(), [q](double val) { return val > q; }); std::cout << "总元素数: " << ba::count(acc) << "\n"; std::cout << "0.99分位数: " << q << "\n"; std::cout << "大于q的元素数: " << count_gt_q << "\n"; return 0; }
方法2:自定义带参数的提取器
如果希望用更贴合Boost Accumulators风格的方式实现,可以自定义一个接受阈值参数的提取器,本质也是基于存储的样本进行统计。
示例代码
#include <boost/accumulators/accumulators.hpp> #include <boost/accumulators/statistics.hpp> #include <boost/accumulators/statistics/ordered_sample.hpp> #include <boost/accumulators/framework/extractor.hpp> #include <boost/range/algorithm/count_if.hpp> #include <iostream> namespace ba = boost::accumulators; // 定义自定义标签 namespace my_tags { struct count_greater_than {}; } // 实现提取器逻辑 namespace boost::accumulators { template<typename AccumulatorSet> std::size_t extract(my_tags::count_greater_than, const AccumulatorSet& acc, double threshold) { const auto& samples = extract(tag::ordered_sample{}, acc); return boost::count_if(samples, [threshold](double val) { return val > threshold; }); } } int main() { Accumulator acc(ba::quantile_probabilities = std::vector<double>{0.99}); // 模拟添加数据 for (int i = 0; i < 1000; ++i) { acc(static_cast<double>(i)); } double q = ba::quantile(acc, ba::quantile_probability = 0.99); // 使用自定义提取器统计大于q的元素数 std::size_t count_gt_q = ba::extract(my_tags::count_greater_than{}, acc, q); std::cout << "大于q的元素数: " << count_gt_q << "\n"; return 0; }
关键注意事项
- 如果使用
p_square_quantile、adaptive_quantile等不存储原始样本的分位数策略,无法事后统计大于q的元素数量,因为accumulator仅保留计算分位数的必要状态,没有原始数据。这种场景下只能在累加元素时实时维护计数器,但无法适配动态变化的q值。 ordered_sample会存储所有样本,适合小数据量场景;如果数据量极大,建议在累加时提前跟踪目标分位数对应的计数,或改用其他统计库(如Boost Histogram)。
内容的提问来源于stack exchange,提问作者SNJ
相关产品推荐
相关产品推荐

