能否使用memcpy反序列化字节?含基本类型的类能否通过memcpy实现序列化反序列化?
Let’s break down your questions one by one, with practical context for C/C++ development:
1. Can you use memcpy to deserialize bytes?
Short answer: It works in very specific, controlled scenarios, but it’s not a safe or portable general-purpose deserialization solution.
memcpy is a low-level function that copies raw bytes from one memory location to another—it has no understanding of the data’s structure or meaning. For it to work for deserialization:
- The source bytes must exactly match the in-memory layout of the target type (including padding, alignment, and byte order).
- The target type must be a POD (Plain Old Data) type (no non-trivial constructors/destructors, no virtual functions, no non-POD members like std::string).
If either condition isn’t met, using memcpy leads to undefined behavior (think corrupted data, crashes, or weird bugs that are hard to track down).
2. Can memcpy serialize/deserialize a class MyClass with basic types?
Your sample code might work if you’re running the same program on the same system, but it’s risky for real-world use. Let’s break down the tradeoffs:
When it might work:
If MyClass is strictly POD (only contains int, char, float, etc., no custom methods or virtual functions) and you’re serializing/deserializing on the same system with the same compiler, build flags, and architecture, memcpy will copy the bytes correctly. The compiler’s padding bytes (added for alignment) will be preserved, and the in-memory layout will match.
The critical problems (why this is not portable):
- Padding bytes: Compilers insert padding between class members to align data with the CPU’s memory access rules. These padding bytes have undefined values. When you serialize, you’re writing these arbitrary bytes to the file. If you deserialize on a system with different padding rules (e.g., different CPU architecture, compiler, or optimization flags), the
MyClasslayout won’t match, leading to corrupted member values. - Endianness: Integer types are stored in memory with a specific byte order (endianness). For example, x86 systems use little-endian, while some embedded systems use big-endian. If you serialize on a little-endian system and deserialize on a big-endian system, integer values will be reversed (e.g.,
0x1234becomes0x3412).memcpydoesn’t handle endian conversion at all. - Alignment risks: While your sample code avoids direct casting (which would be a problem), if you ever tried to cast the
byteDumparray to aMyClass*, you’d hit undefined behavior becausechararrays don’t guarantee alignment for larger types likeint. Even withmemcpy, if the destinationMyClassinstance isn’t properly aligned (though your code avoids this by declaringMyClass toDeserialize;directly), you could run into crashes.
A better, portable alternative:
Instead of relying on raw memory copies, explicitly serialize each member of MyClass. This lets you control endianness and avoid padding issues. For example:
// Serialization (writes to a file/stream with fixed endianness) void serializeMyClass(const MyClass& obj, FILE* file) { // Convert integers to big-endian (network byte order) for portability uint32_t serializedInt = htobe32(obj.myInt); uint16_t serializedShort = htobe16(obj.myShort); fwrite(&serializedInt, sizeof(serializedInt), 1, file); fwrite(&serializedShort, sizeof(serializedShort), 1, file); fwrite(obj.myCharArray, sizeof(obj.myCharArray), 1, file); } // Deserialization (reads from file/stream and converts back to host endianness) void deserializeMyClass(MyClass& obj, FILE* file) { uint32_t serializedInt; uint16_t serializedShort; fread(&serializedInt, sizeof(serializedInt), 1, file); fread(&serializedShort, sizeof(serializedShort), 1, file); fread(obj.myCharArray, sizeof(obj.myCharArray), 1, file); obj.myInt = be32toh(serializedInt); obj.myShort = be16toh(serializedShort); }
This approach works across different systems and compilers because it doesn’t depend on the class’s in-memory layout.
内容的提问来源于stack exchange,提问作者user99999991

