如何在std::vector等顺序容器中存储三种不同编码类型?
可行的实现方案
针对你遇到的三种编码类型存储需求,以下是几种更简洁或内存高效的实现方式:
1. 手动实现紧凑Tagged Union
既然担心带kind成员的结构体内存浪费,可以手动设计内存紧凑的标记联合,将标记与部分数据复用存储:
#include <cassert> #include <utility> struct EncodedData { // header低2位标记类型:0=Pair,1=Bool,2=TinyInt // 类型为TinyInt时,header的2-3位存储0-3的数值 uint8_t header; union { std::pair<int8_t, int8_t> pair; bool b; }; // 构造函数 EncodedData(int8_t a, int8_t b) : header(0) { pair = {a, b}; } EncodedData(bool val) : header(1) { this->b = val; } EncodedData(uint8_t tiny_val) : header(2 | (tiny_val << 2)) { assert(tiny_val <= 3 && "Tiny value must be 0-3"); } // 类型判断与数据获取 enum class Type { Pair, Bool, TinyInt }; Type get_type() const { return static_cast<Type>(header & 0b11); } std::pair<int8_t, int8_t> get_pair() const { assert(get_type() == Type::Pair); return pair; } bool get_bool() const { assert(get_type() == Type::Bool); return b; } uint8_t get_tiny_int() const { assert(get_type() == Type::TinyInt); return (header >> 2) & 0b11; } }; // 使用示例 std::vector<EncodedData> vec; vec.emplace_back(10, -5); vec.emplace_back(true); vec.emplace_back(3);
优点:内存开销极小(结构体大小为3字节);类型判断直接,无额外运行时开销。
缺点:需手动维护构造和访问逻辑,要注意断言或错误检查避免越界访问。
2. 用std::variant配合编译期访问简化代码
你提到std::variant需要大量类型检查,但可以通过std::visit和编译期判断避免重复手动类型判断,让代码更简洁:
#include <variant> #include <vector> #include <utility> // 定义变体类型 using EncodedVariant = std::variant<std::pair<int8_t, int8_t>, bool, uint8_t>; std::vector<EncodedVariant> vec; // 统一处理逻辑的访问器 auto process_data = [](const auto& data) { using T = std::decay_t<decltype(data)>; if constexpr (std::is_same_v<T, std::pair<int8_t, int8_t>>) { // 处理双int8_t类型,例如: // handle_pair(data.first, data.second); } else if constexpr (std::is_same_v<T, bool>) { // 处理布尔类型 // handle_bool(data); } else if constexpr (std::is_same_v<T, uint8_t>) { // 处理0-3的小整数,可加断言限制范围 assert(data <= 3); // handle_tiny_int(data); } }; // 遍历容器处理数据 for (const auto& elem : vec) { std::visit(process_data, elem); }
优点:基于标准库,无需手动管理内存;constexpr if在编译期完成类型分支,运行期无额外开销;代码可读性高,无需维护复杂的联合逻辑。
缺点:变体类型大小等于最大成员大小(此处为2字节)加少量标记开销,内存开销略高于手动联合,但远低于带虚函数的多态类型。
3. 紧凑字节流存储(极致内存优化)
如果追求极致内存利用率,可直接用std::vector<uint8_t>存储序列化后的字节流,手动处理编码和解码:
#include <vector> #include <cassert> // 编码函数 void encode_pair(std::vector<uint8_t>& buf, int8_t a, int8_t b) { buf.push_back(0); // 标记:0=Pair buf.push_back(static_cast<uint8_t>(a)); buf.push_back(static_cast<uint8_t>(b)); } void encode_bool(std::vector<uint8_t>& buf, bool val) { buf.push_back(1); // 标记:1=Bool buf.push_back(static_cast<uint8_t>(val)); } void encode_tiny_int(std::vector<uint8_t>& buf, uint8_t val) { assert(val <= 3); buf.push_back(2 | (val << 2)); // 标记+值:低2位是标记2,高2位存数值 } // 解码函数(需维护当前偏移量) size_t decode(const std::vector<uint8_t>& buf, size_t offset) { uint8_t header = buf[offset]; switch (header & 0b11) { case 0: { // Pair int8_t a = static_cast<int8_t>(buf[offset+1]); int8_t b = static_cast<int8_t>(buf[offset+2]); // 处理a和b return offset + 3; } case 1: { // Bool bool val = static_cast<bool>(buf[offset+1]); // 处理val return offset + 2; } case 2: { // TinyInt uint8_t val = (header >> 2) & 0b11; // 处理val return offset + 1; } default: assert(false && "Invalid header"); return offset; } }
优点:内存利用率最高,无任何额外结构体或变体开销;适合对内存要求极高的场景。
缺点:需手动管理字节流偏移和编码逻辑,易出现越界或解码错误;代码可读性差,维护成本高。
内容的提问来源于stack exchange,提问作者24n8
相关产品推荐
相关产品推荐

