如何在C++中用vector<bool>或Boost::dynamic_bitset获取位图实际字节用于压缩?
我正在开发一款需要处理位图的压缩器,此处的位图指以位而非字节存储布尔值的数组——例如"10010010"这8位数据可存入std::vector<bool>,该容器会将其优化为仅占1字节,而非未优化时的8字节。我希望将内存中的这种优化特性延续至向压缩器传递数据或写入文件时。
可选容器说明
std::bitset需要编译期确定数组大小,不符合动态需求;Boost::dynamic_bitset支持动态扩容,是合适的备选,但我不清楚如何在以下场景中正确使用它。
压缩/解压缩函数声明
现有压缩与解压缩函数的声明如下:
char *compress(const char *data, size_t dataLength, size_t &outSize); char *decompress(const char *data, size_t &compressedSize);
原写法的问题
若要压缩std::vector<bool> bits,我原本的写法是:
// vector<bool> bits is defined and initialized earlier int compressed_size; char* compressed_data = compress(reinterpret_cast<const char*>(&bits), bits.size(), compressed_size);
但这种写法将vector<bool>视为普通布尔数组,会占用更多存储空间。例如100位的数据,bits.size()返回100,若用ZSTD算法压缩,压缩后体积会远大于按13字节(8位/字节)存储的情况。vector<bool>仅在内存中优化存储,无法在写入文件或传递给压缩器时保持该特性。
已尝试的方法及痛点
方法1:使用Boost::dynamic_bitset的公共接口
我查阅了boost::dynamic_bitset<>的源码,发现其私有成员m_bits是存储位数组的vector,每个块为无符号整数。可通过boost::to_block_range()将其拷贝出来实现压缩:
// The definition in dynamic_bitset.hpp template <typename Block, typename Allocator, typename BlockOutputIterator> inline void to_block_range(const dynamic_bitset<Block, Allocator>& b, BlockOutputIterator result) { // note how this copies *all* bits, including the // unused ones in the last block (which are zero) std::copy(b.m_bits.begin(), b.m_bits.end(), result); } // I can use this function to get the vector // boost::dynamic_bitset<> bitset is defined earlier std::vector<boost::dynamic_bitset<>::block_type> filterBlocks(bitset.num_blocks()); // This function copies the internal array out boost::to_block_range(bitset, filterBlocks.begin()); char *compressed_bitset = zstdCompressor.compress(reinterpret_cast<const char *>(&filterBlocks), bitset.num_blocks() * sizeof(boost::dynamic_bitset<>::block_type), compressed_size);
这种方法仅使用公共接口,能正常工作,但会拷贝整个位数组。而我的场景中位数组可能极大,且压缩时无需修改数组,因此不想进行拷贝操作。
方法2:不规范的私有成员访问
我尝试利用to_block_range()作为dynamic_bitset的友元函数可访问m_bits的特性,自定义了该函数的特化版本,试图直接获取内部存储而避免拷贝:
#include <tuple> namespace boost { template <> inline void to_block_range(const dynamic_bitset<Block, Allocator>& b, std::tuple<std::vector<boost::dynamic_bitset<>::block_type>&, size_t&>param) { std::get<0>(param) = b.m_bits; std::get<1>(param) = b.num_blocks() * sizeof(boost::dynamic_bitset<>::block_type); } } std::vector<boost::dynamic_bitset<>::block_type> filterBlocks; size_t dataLength; // This line will call my customized function instead of the original one. Since filterBlocks is a reference, there should be no copy. boost::to_block_range(bitset, std::make_tuple(std::ref(filterBlocks), std::ref(dataLength))); char *compressed_bitset = zstdCompressor.compress(reinterpret_cast<const char *>(&filterBlocks), dataLength, compressed_size);
这种方法通过不规范的方式访问了私有成员m_bits,我不确定是否安全。若Boost::dynamic_bitset内部的块是连续存储的,该方法应该能提升效率,但依赖Boost的内部实现细节,存在兼容性风险。
内容的提问来源于stack exchange,提问作者Yuanjian Liu

