You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C++反序列化:如何告知编译器对象已存在且无需调用构造函数?

问题:C++二进制反序列化中的对象合法性与std::launder的应用

我尝试反序列化文件中的二进制数据,该数据是对象A后跟对象B的有效表示(本问题忽略内存填充)。以下是演示代码,同时需要指出代码中的未定义行为(UB):

#include <iostream>
#include <array>
#include <cstdint>
#include <tuple>

using namespace std;

struct B;

struct A
{
    int some_data = 0x4;
    B* ptr;
};

struct B
{
    int some_int = 0x10;
    long long some_big_int = 0x80;
};

struct ABContainer
{
    A a;
    B b;
};

array<char, 32> serialize(ABContainer* c)
{
    A* a = &c->a;
    B* b = &c->b;

    array<char, 32> r{};
    char* d = r.data();
    int* as_int = reinterpret_cast<int*>(d);
    uintptr_t* as_ptr = reinterpret_cast<uintptr_t*>(d);
    int64_t* as_int64 = reinterpret_cast<int64_t*>(d);

    // Stamp A values in r
    *as_int = a->some_data;
    // Store only the offset, not the full address
    as_ptr[1] = reinterpret_cast<uintptr_t>(a->ptr) - reinterpret_cast<uintptr_t>(a);

    // Stamp B values in r
    as_int[4] = b->some_int;
    as_int64[3] = b->some_big_int;

    return r;
}

std::tuple<A*, B*> deserialize(array<char, 32>& source)
{
    char* d = source.data();
    
    uintptr_t* as_ptr = reinterpret_cast<uintptr_t*>(d);

    // Adjust ptr to point again at a valid address
    as_ptr[1] += reinterpret_cast<uintptr_t>(d);

    A* a = reinterpret_cast<A*>(d);
    B* b = reinterpret_cast<B*>(d + 16); // Hard code +16 to skip A

    return { a, b };
}

int main()
{
    // Use a container to remove stack alignment for the demo
    ABContainer c;
    c.a.ptr = &c.b;

    // serialized is the buffer that would have been filled by reading a file
    array<char, 32> serialized = serialize(&c);
    auto [a_obj, b_obj] = deserialize(serialized);
    
    cout << "a->some_data: " << a_obj->some_data << endl;
    cout << "a->ptr->some_int: " << a_obj->ptr->some_int << endl;
    cout << "a->ptr->some_big_int: " << a_obj->ptr->some_big_int << endl;
    cout << "b->some_int: " << b_obj->some_int << endl;
    cout << "b->some_big_int: " << b_obj->some_big_int << endl;
}

原代码中的未定义行为

  • 直接通过reinterpret_cast将char*转换为A*/B*并解引用:char数组的内存中从未构造过A或B对象,编译器无法识别该地址处存在合法对象,任何解引用操作都属于未定义行为。
  • 通过uintptr_t*直接修改内存内容来调整A::ptr:这相当于绕过对象构造流程直接篡改内存,违反了C++对象生命周期的规则,同样属于未定义行为。

std::launder的适用性

std::launder的核心作用是告知编译器:某地址处已经存在一个合法构造的对象,需要返回指向该对象的有效指针。但它不能在无对象的内存中凭空生成合法对象——原代码中只是把二进制数据拷贝到了char数组,对应位置根本没有构造过A和B对象,因此std::launder无法解决问题,不适用此场景。

符合C++标准的修改方案

要合法完成反序列化,必须在char数组的对应内存位置显式构造对象,可以通过placement new实现,同时序列化时应直接读写对象成员而非依赖内存布局拷贝:

修改后的序列化函数

array<char, 32> serialize(ABContainer* c)
{
    array<char, 32> r{};
    char* d = r.data();

    // 写入A的成员值
    *reinterpret_cast<int*>(d) = c->a.some_data;
    uintptr_t offset = reinterpret_cast<uintptr_t>(c->a.ptr) - reinterpret_cast<uintptr_t>(&c->a);
    *reinterpret_cast<uintptr_t*>(d + sizeof(int)) = offset;

    // 写入B的成员值
    char* b_dest = d + 16;
    *reinterpret_cast<int*>(b_dest) = c->b.some_int;
    *reinterpret_cast<int64_t*>(b_dest + sizeof(int)) = c->b.some_big_int;

    return r;
}

修改后的反序列化函数

std::tuple<A*, B*> deserialize(array<char, 32>& source)
{
    char* d = source.data();
    
    // 在d指向的内存构造A对象
    A* a = new (d) A;
    // 从序列化数据中读取成员值
    a->some_data = *reinterpret_cast<int*>(d);
    uintptr_t offset = *reinterpret_cast<uintptr_t*>(d + sizeof(int));
    a->ptr = reinterpret_cast<B*>(reinterpret_cast<uintptr_t>(d) + offset);

    // 在d+16的位置构造B对象
    char* b_addr = d + 16;
    B* b = new (b_addr) B;
    // 从序列化数据中读取成员值
    b->some_int = *reinterpret_cast<int*>(b_addr);
    b->some_big_int = *reinterpret_cast<int64_t*>(b_addr + sizeof(int));

    return {a, b};
}

补充:对象的销毁

由于使用placement new构造的对象不会自动调用析构函数,若对象拥有非平凡析构函数(比如包含动态内存、智能指针等),需要手动调用析构函数:

// 在main函数末尾添加
a_obj->~A();
b_obj->~B();

内容的提问来源于stack exchange,提问作者Sproulx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 23:48:17