内存流序列化反序列化内存占用过高原因及优化建议咨询
Let's break down why you're seeing those dramatic memory spikes and how to fix them step by step.
Why the Memory Spikes Happen
1. BinaryFormatter Causes Massive Temporary Overhead
BinaryFormatter is notoriously inefficient for large objects like DataTables. When you call f.Serialize(ms, dc), it doesn't just write raw bytes to the MemoryStream—it creates a huge number of intermediate objects to traverse and serialize the entire object graph of the DataTable.
Your 1-million-row DataTable with three string columns means BinaryFormatter has to serialize every detail of the table's structure (columns, constraints, row states) plus each row's data. This generates tons of temporary objects that live in the managed heap, leading to memory usage 8-10x larger than the final serialized byte array (hence the 3GB spike even though your payload is only ~338MB).
2. Large Object Heap (LOH) Retains Memory After Nulling Variables
When you set dc = null and byteSD = null, you mark those objects for garbage collection—but .NET's GC doesn't clean up immediately, especially for large objects (anything over 85KB, like your 338MB byte array). These live in the Large Object Heap (LOH), which is collected far less frequently than the small object heap.
Additionally, the LOH isn't compacted by default. Even after garbage collection, the memory might remain reserved (showing as 1.6GB in your process) until the OS reclaims it or the GC faces enough pressure to perform a full LOH cleanup.
Optimization Strategies to Fix Memory Issues
1. Replace BinaryFormatter with a Modern, Efficient Serializer
BinaryFormatter is obsolete and inefficient—swap it for a serializer designed for performance. Here are the best options:
Option A: Use Protobuf-net (Fast, Low Memory Overhead)
Protobuf-net generates far fewer temporary objects and produces smaller payloads. To use it:
- First, replace the heavy DataTable with a strongly-typed list (POCOs serialize much more efficiently):
[Serializable] [ProtoContract] public class DataContainer { [ProtoMember(1)] public string TableName { get; set; } [ProtoMember(2)] public List<TestRow> Rows { get; set; } } [ProtoContract] public class TestRow { [ProtoMember(1)] public string Column1 { get; set; } [ProtoMember(2)] public string Column2 { get; set; } [ProtoMember(3)] public string Column3 { get; set; } } - Update your
CreateTablemethod to populate the list instead of a DataTable (pre-allocate capacity to avoid resizing overhead):private void CreateTable() { var rows = new List<TestRow>(1000000); string one = Guid.NewGuid().ToString(); string two = Guid.NewGuid().ToString(); string three = Guid.NewGuid().ToString(); for (int i = 0; i < 1000000; i++) { rows.Add(new TestRow { Column1 = one, Column2 = two, Column3 = three }); } dc.Rows = rows; } - Serialize/Deserialize with Protobuf-net:
private void SerialiseObj() { using (MemoryStream ms = new MemoryStream()) { Serializer.Serialize(ms, dc); byteSD = ms.ToArray(); } } private void DeserialiseObj() { using (MemoryStream ms = new MemoryStream(byteSD)) { var _dc = Serializer.Deserialize<DataContainer>(ms); // Use _dc then let it go out of scope } }
This will cut your serialization memory usage drastically—you'll likely see spikes under 500MB instead of 3GB.
Option B: Use MessagePack (Extremely Compact, Low Memory)
MessagePack is another great choice for high-performance serialization, especially for large datasets. It uses a compact binary format and minimizes temporary object creation.
2. Optimize MemoryStream Usage
Your original SerialiseObj method has a redundant allocation:
byteSD = new byte[ms.Length]; // Unnecessary allocation byteSD = ms.ToArray(); // Creates a new array anyway
Remove the redundant line—just use byteSD = ms.ToArray();. Also, never call ms.Dispose() inside a using block; the using statement handles disposal automatically.
3. Use ArrayPool for Large Byte Arrays
Instead of allocating new large arrays, use ArrayPool<byte> to reuse existing buffers and reduce LOH fragmentation:
private void SerialiseObj() { using (MemoryStream ms = new MemoryStream()) { Serializer.Serialize(ms, dc); int length = (int)ms.Length; byteSD = ArrayPool<byte>.Shared.Rent(length); ms.Position = 0; ms.Read(byteSD, 0, length); } } // In DeserialiseObj, after use: ArrayPool<byte>.Shared.Return(byteSD); byteSD = null;
This avoids creating new large objects in the LOH, reducing memory retention.
4. Trigger Garbage Collection (Carefully)
If you need to reclaim memory immediately (e.g., in a test scenario), you can force a full GC collection. Note: This is not recommended for production code unless absolutely necessary, as it disrupts the GC's automatic optimization:
dc = null; byteSD = null; GC.Collect(2, GCCollectionMode.Forced, true, true); GC.WaitForPendingFinalizers(); GC.Collect(2, GCCollectionMode.Forced, true, true);
This triggers a full collection of all heaps, including the LOH, and should bring memory closer to the initial 17MB.
5. Consider Chunked Serialization/Transmission
If your DataTable is too large to handle in memory at once, split it into smaller chunks (e.g., 10,000 rows per chunk), serialize each chunk separately, and transmit them one by one. This keeps memory usage low throughout the process.
Why Your Original Code Didn't Release Memory
- The LOH doesn't get collected as often as the small object heap. Even after nulling variables, the GC won't collect the LOH until it detects significant memory pressure.
- BinaryFormatter leaves behind many temporary objects in the heap that take time to clean up. Modern serializers avoid this by using more efficient serialization logic.
内容的提问来源于stack exchange,提问作者Gizazas

