You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在std::vector等顺序容器中存储三种不同编码类型?

可行的实现方案

针对你遇到的三种编码类型存储需求,以下是几种更简洁或内存高效的实现方式:

1. 手动实现紧凑Tagged Union

既然担心带kind成员的结构体内存浪费,可以手动设计内存紧凑的标记联合,将标记与部分数据复用存储:

#include <cassert>
#include <utility>

struct EncodedData {
    // header低2位标记类型:0=Pair,1=Bool,2=TinyInt
    // 类型为TinyInt时,header的2-3位存储0-3的数值
    uint8_t header;
    union {
        std::pair<int8_t, int8_t> pair;
        bool b;
    };

    // 构造函数
    EncodedData(int8_t a, int8_t b) : header(0) {
        pair = {a, b};
    }
    EncodedData(bool val) : header(1) {
        this->b = val;
    }
    EncodedData(uint8_t tiny_val) : header(2 | (tiny_val << 2)) {
        assert(tiny_val <= 3 && "Tiny value must be 0-3");
    }

    // 类型判断与数据获取
    enum class Type { Pair, Bool, TinyInt };
    Type get_type() const {
        return static_cast<Type>(header & 0b11);
    }

    std::pair<int8_t, int8_t> get_pair() const {
        assert(get_type() == Type::Pair);
        return pair;
    }
    bool get_bool() const {
        assert(get_type() == Type::Bool);
        return b;
    }
    uint8_t get_tiny_int() const {
        assert(get_type() == Type::TinyInt);
        return (header >> 2) & 0b11;
    }
};

// 使用示例
std::vector<EncodedData> vec;
vec.emplace_back(10, -5);
vec.emplace_back(true);
vec.emplace_back(3);

优点:内存开销极小(结构体大小为3字节);类型判断直接,无额外运行时开销。
缺点:需手动维护构造和访问逻辑,要注意断言或错误检查避免越界访问。

2. 用std::variant配合编译期访问简化代码

你提到std::variant需要大量类型检查,但可以通过std::visit和编译期判断避免重复手动类型判断,让代码更简洁:

#include <variant>
#include <vector>
#include <utility>

// 定义变体类型
using EncodedVariant = std::variant<std::pair<int8_t, int8_t>, bool, uint8_t>;
std::vector<EncodedVariant> vec;

// 统一处理逻辑的访问器
auto process_data = [](const auto& data) {
    using T = std::decay_t<decltype(data)>;
    if constexpr (std::is_same_v<T, std::pair<int8_t, int8_t>>) {
        // 处理双int8_t类型,例如:
        // handle_pair(data.first, data.second);
    } else if constexpr (std::is_same_v<T, bool>) {
        // 处理布尔类型
        // handle_bool(data);
    } else if constexpr (std::is_same_v<T, uint8_t>) {
        // 处理0-3的小整数,可加断言限制范围
        assert(data <= 3);
        // handle_tiny_int(data);
    }
};

// 遍历容器处理数据
for (const auto& elem : vec) {
    std::visit(process_data, elem);
}

优点:基于标准库,无需手动管理内存;constexpr if在编译期完成类型分支,运行期无额外开销;代码可读性高,无需维护复杂的联合逻辑。
缺点:变体类型大小等于最大成员大小(此处为2字节)加少量标记开销,内存开销略高于手动联合,但远低于带虚函数的多态类型。

3. 紧凑字节流存储(极致内存优化)

如果追求极致内存利用率,可直接用std::vector<uint8_t>存储序列化后的字节流,手动处理编码和解码:

#include <vector>
#include <cassert>

// 编码函数
void encode_pair(std::vector<uint8_t>& buf, int8_t a, int8_t b) {
    buf.push_back(0); // 标记:0=Pair
    buf.push_back(static_cast<uint8_t>(a));
    buf.push_back(static_cast<uint8_t>(b));
}
void encode_bool(std::vector<uint8_t>& buf, bool val) {
    buf.push_back(1); // 标记:1=Bool
    buf.push_back(static_cast<uint8_t>(val));
}
void encode_tiny_int(std::vector<uint8_t>& buf, uint8_t val) {
    assert(val <= 3);
    buf.push_back(2 | (val << 2)); // 标记+值:低2位是标记2,高2位存数值
}

// 解码函数(需维护当前偏移量)
size_t decode(const std::vector<uint8_t>& buf, size_t offset) {
    uint8_t header = buf[offset];
    switch (header & 0b11) {
        case 0: { // Pair
            int8_t a = static_cast<int8_t>(buf[offset+1]);
            int8_t b = static_cast<int8_t>(buf[offset+2]);
            // 处理a和b
            return offset + 3;
        }
        case 1: { // Bool
            bool val = static_cast<bool>(buf[offset+1]);
            // 处理val
            return offset + 2;
        }
        case 2: { // TinyInt
            uint8_t val = (header >> 2) & 0b11;
            // 处理val
            return offset + 1;
        }
        default:
            assert(false && "Invalid header");
            return offset;
    }
}

优点:内存利用率最高,无任何额外结构体或变体开销;适合对内存要求极高的场景。
缺点:需手动管理字节流偏移和编码逻辑,易出现越界或解码错误;代码可读性差,维护成本高。


内容的提问来源于stack exchange,提问作者24n8

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 04:05:54