迁移数据库50万共700GB文件到MinIO时出现StackOverflow异常无法定位原因
问题背景
我正在将总量约50万份、总大小700GB的文件从数据库迁移至MinIO对象存储,使用控制台程序实现,但程序存在内存泄漏问题,运行后会抛出StackOverflow异常,无法定位具体原因。
原始核心逻辑代码
var watch = System.Diagnostics.Stopwatch.StartNew(); Console.WriteLine("\n"); Console.WriteLine("开始迁移"); Console.WriteLine("正在获取元数据"); var attachments = await dbContext.Attachments.ToListAsync(); Console.WriteLine("元数据获取完成"); Console.WriteLine($"附件总数量: {attachments.Count}"); Console.WriteLine("\n"); int total = 0; int completed = 0; int faulted = 0; bool fileAlreadyMigrated = false; using (StreamWriter sw = new StreamWriter("migrationLogs.txt")) { foreach (var attachment in attachments) { total++; Console.WriteLine($"开始迁移ID为 {attachment.Id} 的附件,进度:{total}/{attachments.Count} ({((double)total / (double)attachments.Count) * 100: 00.##}%)"); sw.WriteLine($"开始迁移ID为 {attachment.Id} 的附件,进度:{total}/{attachments.Count} ({((double)total / (double)attachments.Count) * 100: 00.##}%)"); attachment.Migrated = false; attachment.Faulted = false; try { fileAlreadyMigrated = false; using (var migratedData = new StreamReader("migratedData.txt")) { while (migratedData.Peek() != -1) { if (Guid.Parse(migratedData.ReadLine()) == attachment.Id) fileAlreadyMigrated = true; } } if (!fileAlreadyMigrated) { var binary = await dbContext.AttachmentBinaries.FirstOrDefaultAsync(b => b.BinaryId == attachment.BinaryId); attachment.FileBinary = binary.FileInBinary; var res = await storageService.UploadBinary(attachment); if (res == System.Net.HttpStatusCode.OK) { Console.WriteLine($"ID为 {attachment.Id} 的附件迁移成功"); sw.WriteLine($"ID为 {attachment.Id} 的附件迁移成功"); attachment.Migrated = true; completed++; File.AppendAllText("migratedData.txt", attachment.Id + Environment.NewLine); } else { Console.WriteLine($"ID为 {attachment.Id} 的附件迁移失败\n错误信息:{res.ToString()}"); sw.WriteLine($"ID为 {attachment.Id} 的附件迁移失败\n错误信息:{res.ToString()}"); attachment.Faulted = true; faulted++; } } else { continue; } } catch (Exception e) { Console.WriteLine($"ID为 {attachment.Id} 的附件迁移失败\n错误信息:{e.Message} {e.InnerException?.Message}"); sw.WriteLine($"ID为 {attachment.Id} 的附件迁移失败\n错误信息:{e.Message} {e.InnerException?.Message}"); attachment.Faulted = true; faulted++; } Console.WriteLine("\n"); Thread.Sleep(1000); } watch.Stop(); var timeStamp = TimeSpan.FromMilliseconds(watch.ElapsedMilliseconds); Console.WriteLine($"迁移完成,耗时:{timeStamp.Hours:###00}:{timeStamp.Minutes:00}:{timeStamp.Seconds:00},{timeStamp.Milliseconds}"); Console.WriteLine($"总处理附件数:{total}"); Console.WriteLine($"迁移成功数:{completed}"); Console.WriteLine($"迁移失败数:{faulted}"); Console.WriteLine($"成功率:{((double)completed / (double)total) * 100:00.##}%"); Console.WriteLine("\n"); sw.WriteLine($"迁移完成,耗时:{timeStamp.Hours:###00}:{timeStamp.Minutes:00}:{timeStamp.Seconds:00},{timeStamp.Milliseconds}"); sw.WriteLine($"总处理附件数:{total}"); sw.WriteLine($"迁移成功数:{completed}"); sw.WriteLine($"迁移失败数:{faulted}"); sw.WriteLine($"成功率:{((double)completed / (double)total) * 100:00.##}%"); sw.WriteLine("\n"); var userInput = string.Empty; do { Console.WriteLine("是否导出迁移失败的附件ID列表?(Y/N)"); userInput = Console.ReadLine(); } while (userInput.ToLower() != "y" && userInput.ToLower() != "n"); if (userInput.ToLower() == "y") { foreach (var attachment in attachments.Where(a => a.Faulted)) { Console.WriteLine(attachment.Id); } } } Console.WriteLine("\n");
UploadBinary方法实现
public async Task<HttpStatusCode> UploadBinary(Attachment dto) { var response = await httpService.SendPostRequestAsync<Attachment>(storageUrl, dto); return response; }
已定位问题与优化过程
最初猜测问题和dbContext生命周期过长、未使用AsNoTracking()有关,后续定位到第一个核心问题:
attachment.FileBinary = binary.FileInBinary;
遍历attachment对象时,每一步都会将文件二进制内容赋值到attachment对象中,而所有attachment对象都存放在全局加载的attachments列表里,导致所有二进制数据都被列表引用无法释放,内存占用持续升高。
将二进制内容存入临时变量后,内存上涨速度变慢,但每轮迭代内存仍会升高。修改为如下代码后问题解决:
var fileInBinary = dbContext.AttachmentBinaries.AsNoTracking().FirstOrDefault(b => b.BinaryId == attachment.BinaryId).FileInBinary; var res = storageService.UploadBinary(attachment, fileInBinary); Array.Clear(fileInBinary, 0, fileInBinary.Length);
遗留疑问
按理说每轮迭代结束后临时变量没有引用了,GC应该会自动回收,为什么还需要显式清空数组呢?
问题解答
临时变量未被GC及时回收的原因
- GC回收时机不确定:GC并不会在变量离开作用域后立刻执行回收,而是会在内存占用达到设定阈值时才触发回收。迁移场景下每轮都会加载MB级的二进制数组,回收速度赶不上内存分配速度,就会出现内存持续上涨的情况。
- 大对象堆(LOH)特性:大于85000字节的对象会被分配到大对象堆,LOH的回收频率远低于普通新生代堆,且默认不会进行碎片整理,即使没有引用的大对象也会占用内存空间,很容易导致内存泄漏和内存碎片化问题。
- 隐式引用残留:http服务发送请求时,可能会将二进制数组的引用存放在请求缓冲区、异步上下文等位置,这些隐式引用的释放存在延迟,会导致数组无法在迭代结束后立刻进入可回收状态。
额外优化建议
- 不要一次性加载所有50万条Attachment元数据到内存,可以分页查询,每次只处理100-500条,处理完后就释放这部分对象的引用,降低全局内存占用。
- 尽量使用流式上传,不要把整个二进制数组加载到内存,直接从数据库读二进制流写入MinIO请求流,彻底避免大数组内存分配。
- 已迁移的ID列表不要每次遍历都读一遍文件,可以提前加载到
HashSet<Guid>中,判断是否已迁移的时间复杂度从O(n)降到O(1),同时避免频繁IO操作。 - dbContext可以每处理一批就重建一次,避免跟踪缓存堆积。
内容的提问来源于stack exchange,提问作者Jamil
相关产品推荐
相关产品推荐

