如何用Boost Multi-index实现GROUP BY COUNT统计功能
使用Boost.MultiIndex对卫星数据分组统计
我正在用boost::multi_index_container处理卫星数据,定义的数据结构如下:
struct SatData { std::string sat_system; std::string band; int index; uint64_t time; double data; SatData(const std::string& ss, const std::string& b, const int& i, const uint64_t& t, const double& d) : sat_system(ss), band(b), index(i), time(t), data(d) {} }; using SatDataset = multi_index_container <SatData, indexed_by < ordered_unique // 给定特定卫星和时间,测量数据唯一 < composite_key < SatData, member<SatData, std::string, &SatData::sat_system>, member<SatData, std::string, &SatData::band>, member<SatData, int, &SatData::index>, member<SatData, uint64_t, &SatData::time> > // composite key > // ordered_unique >; // multi_index_container
我需要按sat_system+band+index分组,统计每组的测量次数。例如有如下数据:
GPS | L1 | 1 | 10 GPS | L1 | 1 | 11 GPS | L1 | 1 | 12 GPS | L2 | 1 | 10 GPS | L2 | 1 | 11 GPS | L2 | 1 | 12 GPS | L2 | 4 | 11 GPS | L2 | 4 | 12 GALILEO | E5b | 2 | 10 GLONASS | G1 | 1 | 10 GLONASS | G1 | 1 | 11 GLONASS | G1 | 2 | 10
期望得到的统计结果:
GPS | L1 | 1 | 3 GPS | L2 | 1 | 2 GPS | L2 | 1 | 1 GPS | L2 | 4 | 2 GALILEO | E5b | 2 | 1 GLONASS | G1 | 1 | 3
注:原示例期望结果中GPS|L2|1的统计可能存在笔误,实际对应数据的统计数应为3。
对应的伪SQL语句:
SELECT sat_system, band, index, COUNT(time) GROUP BY sat_system, band, index
我之前找到过相关问题,但都是针对单个查询场景的,现在需要对全表做分组统计。
解决方案
可以利用Boost.MultiIndex的有序索引特性,通过迭代器遍历实现全表分组统计,提供两种实现方式:
方法1:基于现有索引直接统计
无需修改容器定义,利用现有复合索引的前缀字段(sat_system、band、index)进行分组:
#include <iostream> #include <boost/multi_index_container.hpp> #include <boost/multi_index/ordered_index.hpp> #include <boost/multi_index/member.hpp> #include <boost/multi_index/composite_key.hpp> #include <boost/tuple/tuple.hpp> // SatData和SatDataset定义同前文 int main() { SatDataset dataset; // 插入示例数据 dataset.insert(SatData("GPS", "L1", 1, 10, 0.0)); dataset.insert(SatData("GPS", "L1", 1, 11, 0.0)); dataset.insert(SatData("GPS", "L1", 1, 12, 0.0)); dataset.insert(SatData("GPS", "L2", 1, 10, 0.0)); dataset.insert(SatData("GPS", "L2", 1, 11, 0.0)); dataset.insert(SatData("GPS", "L2", 1, 12, 0.0)); dataset.insert(SatData("GPS", "L2", 4, 11, 0.0)); dataset.insert(SatData("GPS", "L2", 4, 12, 0.0)); dataset.insert(SatData("GALILEO", "E5b", 2, 10, 0.0)); dataset.insert(SatData("GLONASS", "G1", 1, 10, 0.0)); dataset.insert(SatData("GLONASS", "G1", 1, 11, 0.0)); dataset.insert(SatData("GLONASS", "G1", 2, 10, 0.0)); // 获取容器的有序唯一索引 auto& idx = dataset.get<0>(); auto it = idx.begin(); while (it != idx.end()) { // 构造分组键前缀 auto group_key = boost::make_tuple(it->sat_system, it->band, it->index); // 匹配所有同前缀的元素范围 auto range = idx.equal_range(boost::make_tuple(group_key, boost::tuples::open())); // 统计分组内元素数量 std::size_t count = std::distance(range.first, range.second); // 输出统计结果 std::cout << it->sat_system << "\t| " << it->band << "\t| " << it->index << "\t| " << count << "\n"; // 跳转到下一个分组 it = range.second; } return 0; }
方法2:新增专用索引优化性能
如果需要频繁执行分组统计,推荐给容器添加一个仅包含sat_system、band、index的有序非唯一索引,减少索引冗余,提升遍历效率:
// 修改SatDataset定义,新增分组专用索引 using SatDataset = multi_index_container <SatData, indexed_by < ordered_unique // 原唯一索引 < composite_key < SatData, member<SatData, std::string, &SatData::sat_system>, member<SatData, std::string, &SatData::band>, member<SatData, int, &SatData::index>, member<SatData, uint64_t, &SatData::time> > >, ordered_non_unique // 新增分组专用索引 < composite_key < SatData, member<SatData, std::string, &SatData::sat_system>, member<SatData, std::string, &SatData::band>, member<SatData, int, &SatData::index> > > >;
基于新索引的统计代码:
// 获取分组专用索引 auto& group_idx = dataset.get<1>(); auto it = group_idx.begin(); while (it != group_idx.end()) { auto group_key = boost::make_tuple(it->sat_system, it->band, it->index); auto range = group_idx.equal_range(group_key); std::size_t count = std::distance(range.first, range.second); std::cout << it->sat_system << "\t| " << it->band << "\t| " << it->index << "\t| " << count << "\n"; it = range.second; }
两种方法都能实现需求,方法2在高频统计场景下性能更优。
内容的提问来源于stack exchange,提问作者mcamurri
相关产品推荐
相关产品推荐

