现代C++中处理可选类成员:规避静态hack与代码重复的方法
在处理千兆字节级的海量数据时,我通过数组索引访问数据。为提升缓存效率,希望将数组的部分数据与索引一起缓存,加快基于索引的操作速度。
缓存数据量为编译期可配置选项,包含缓存量为0的场景。由于要处理大量索引,缓存量为0时我不希望像std::array那样产生额外的空元素内存开销。
我目前实现了带特化的模板:
using index_t = unsigned int; using lexem_t = unsigned int; template <std::size_t t_arg_cache_line_size> struct lexem_index_with_cache_t { index_t index; std::array<lexem_t, t_arg_cache_line_size> cache_line; constexpr std::size_t cache_line_size() const { return t_arg_cache_line_size; } }; template<> struct lexem_index_with_cache_t<0> { index_t index; static std::array<lexem_t, 0> cache_line; constexpr std::size_t cache_line_size() const { return 0; } }; std::array<lexem_t, 0> lexem_index_with_cache_t<0>::cache_line;
这里的问题是,在0大小的特化中我用了一个“hack”:通过静态成员cache_line提供形式上的访问接口,实际上这个成员是空的且不会被访问。这么做是为了让使用该模板的函数无需特化,比如下面的比较器:
using lexem_index_with_cache = lexem_index_with_cache_t<0>; template <typename T> class seq_forward_comparator_cached { const std::vector<T>& vec; public: seq_forward_comparator_cached(const std::vector<T>& vec) : vec(vec) { } bool operator() (const lexem_index_with_cache& idx1, const lexem_index_with_cache& idx2) { if (idx1.index == idx2.index) { return false; } const auto it1_cache_line = idx1.cache_line; // 无静态hack时这段代码无法编译 const auto it2_cache_line = idx2.cache_line; // 无静态hack时这段代码无法编译 auto res = std::lexicographical_compare_three_way( it1_cache_line.begin(), it1_cache_line.end(), it2_cache_line.begin(), it2_cache_line.end()); if (res == std::strong_ordering::equal) { auto range1 = std::ranges::subrange(vec.begin() + idx1.index + idx1.cache_line_size(), vec.end()); auto range2 = std::ranges::subrange(vec.begin() + idx2.index + idx2.cache_line_size(), vec.end()); return std::ranges::lexicographical_compare(range1, range2); } return res == std::strong_ordering::less; } };
我可以为零大小缓存的情况实现模板特化,但这样会导致大量代码重复,而我有很多类似的函数,不想全部做特化。
请问在现代C++中,有没有规范的方法可以避免这种static hack,同时又能避免代码重复?我不确定是否可以通过依赖类型的条件代码包含来解决,希望尽量不封装cache_line的访问函数,如果这是唯一可行的方案,请提供思路。
方法1:使用constexpr if消除静态hack
在C++17及以后,可利用constexpr if在函数内部根据缓存大小分支处理,无需给0大小特化添加静态cache_line成员,也不用特化整个函数。
首先修改0大小的模板特化,去掉静态成员:
template<> struct lexem_index_with_cache_t<0> { index_t index; constexpr std::size_t cache_line_size() const { return 0; } };
然后修改比较器的operator(),用constexpr if区分缓存大小为0和非0的情况:
template <typename T> class seq_forward_comparator_cached { const std::vector<T>& vec; public: seq_forward_comparator_cached(const std::vector<T>& vec) : vec(vec) { } template <std::size_t CacheSize> bool operator() (const lexem_index_with_cache_t<CacheSize>& idx1, const lexem_index_with_cache_t<CacheSize>& idx2) { if (idx1.index == idx2.index) { return false; } std::strong_ordering res = std::strong_ordering::equal; if constexpr (CacheSize > 0) { res = std::lexicographical_compare_three_way( idx1.cache_line.begin(), idx1.cache_line.end(), idx2.cache_line.begin(), idx2.cache_line.end()); } if (res == std::strong_ordering::equal) { auto range1 = std::ranges::subrange(vec.begin() + idx1.index + idx1.cache_line_size(), vec.end()); auto range2 = std::ranges::subrange(vec.begin() + idx2.index + idx2.cache_line_size(), vec.end()); return std::ranges::lexicographical_compare(range1, range2); } return res == std::strong_ordering::less; } };
当CacheSize为0时,编译器会直接跳过缓存比较的代码分支,既不需要静态hack,也避免了代码重复。
方法2:用类型萃取封装缓存访问(如需保持统一接口)
如果希望保持cache_line的统一访问形式,可通过类型萃取或辅助函数封装,避免直接暴露成员:
首先定义辅助模板:
template <std::size_t CacheSize> struct cache_accessor { static constexpr auto& get(const lexem_index_with_cache_t<CacheSize>& obj) { return obj.cache_line; } }; template <> struct cache_accessor<0> { static constexpr std::array<lexem_t, 0> get(const lexem_index_with_cache_t<0>&) { return {}; } };
然后在比较器中使用这个辅助类:
bool operator() (const lexem_index_with_cache& idx1, const lexem_index_with_cache& idx2) { if (idx1.index == idx2.index) { return false; } const auto it1_cache_line = cache_accessor<0>::get(idx1); const auto it2_cache_line = cache_accessor<0>::get(idx2); auto res = std::lexicographical_compare_three_way( it1_cache_line.begin(), it1_cache_line.end(), it2_cache_line.begin(), it2_cache_line.end()); // 后续逻辑不变 }
这种方式可保持代码一致性,同时消除静态hack,但需要额外封装访问逻辑。
方法3:利用空基类优化(EBO)
另一种思路是将cache_line作为基类,利用空基类优化避免0大小数组的内存开销:
template <std::size_t CacheSize> struct cache_storage : std::array<lexem_t, CacheSize> {}; template <> struct cache_storage<0> {}; // 空基类,无内存开销 template <std::size_t t_arg_cache_line_size> struct lexem_index_with_cache_t : cache_storage<t_arg_cache_line_size> { index_t index; constexpr std::size_t cache_line_size() const { return t_arg_cache_line_size; } // 提供统一的cache_line访问接口 auto& cache_line() { return static_cast<cache_storage<t_arg_cache_line_size>&>(*this); } const auto& cache_line() const { return static_cast<const cache_storage<t_arg_cache_line_size>&>(*this); } };
然后在比较器中调用cache_line()成员函数:
auto res = std::lexicographical_compare_three_way( idx1.cache_line().begin(), idx1.cache_line().end(), idx2.cache_line().begin(), idx2.cache_line().end());
这种方式通过空基类优化避免了0大小的内存开销,同时提供了统一的访问接口,不需要静态hack,也无需特化函数。
内容的提问来源于stack exchange,提问作者Damir Tenishev

