C#大列表拆分:将百万级文件按99条分组断点处理
Handling Large Files with Grouping & Resume-from-Breakpoint in C#
Hey there! As someone who's been where you are (struggling with large file processing as a C# newbie), let's work through this problem step by step. We'll fix your code to handle that 10M-line password file, split it into groups of 99, pick up right where you left off if the program stops, and make it flexible enough for any list-based file.
First, Let's Break Down the Issues in Your Original Code
File.ReadAllLines()loads the entire 10M-line file into memory at once—this will chew up tons of RAM and might even crash your program.- No way to track how many lines you've already processed, so you can't resume after a break.
- No logic to group lines into chunks of 99.
The Solution: Efficient, Resumable Group Processing
Here's a revamped version of your code that fixes all these gaps. It uses memory-friendly line-by-line reading, tracks progress with a simple breakpoint file, and lets you tweak settings for any file:
using System; using System.IO; namespace ListProcessingWithResume { class Program { static void Main(string[] args) { // Configure these for your use case (easy to adapt to any file!) string inputFilePath = @"c:\tmp\10-million-password-list-top-1000000.txt"; int groupSize = 99; string breakpointFilePath = @"c:\tmp\last_processed_line.txt"; // Load progress from breakpoint file if it exists int linesProcessed = 0; if (File.Exists(breakpointFilePath)) { if (int.TryParse(File.ReadAllText(breakpointFilePath), out int savedCount)) { linesProcessed = savedCount; Console.WriteLine($"Resuming from line {linesProcessed + 1}..."); } } int currentGroupCount = 0; // Use StreamReader to read one line at a time (memory-friendly for big files) using (StreamReader reader = new StreamReader(inputFilePath)) { string line; // Skip lines we've already processed for (int i = 0; i < linesProcessed; i++) { reader.ReadLine(); } while ((line = reader.ReadLine()) != null) { linesProcessed++; currentGroupCount++; // Add line to current group (modify this to save/process the group as needed) Console.WriteLine($"Line {linesProcessed}: {line}"); // When we hit the group size, wrap up the group and update progress if (currentGroupCount == groupSize) { Console.WriteLine($"--- Finished group ending at line {linesProcessed} ---"); // Save progress so we can resume later File.WriteAllText(breakpointFilePath, linesProcessed.ToString()); currentGroupCount = 0; // Optional: Add a pause if you need to throttle processing // System.Threading.Thread.Sleep(100); } } // Handle any leftover lines that didn't make a full group if (currentGroupCount > 0) { Console.WriteLine($"--- Finished final partial group with {currentGroupCount} lines (ending at line {linesProcessed}) ---"); File.WriteAllText(breakpointFilePath, linesProcessed.ToString()); } Console.WriteLine("All lines processed successfully!"); // Optional: Delete the breakpoint file if you don't need it anymore // File.Delete(breakpointFilePath); } } } }
Key Features Explained (For Newbies!)
- Memory Efficiency:
StreamReaderreads one line at a time instead of loading the whole file, so even 10M lines won't bog down your system. - Breakpoint Resumption: The
breakpointFilePathstores the last line number processed. If the program crashes or you close it, it picks up right where it left off next time. - Flexible Grouping: Just change the
groupSizevariable to split into any number of lines per group (not just 99). - Universal Use: Swap out
inputFilePathto process any text file with one item per line—no need to rewrite the whole code.
How to Use It
- Set your input file path, group size, and breakpoint file path at the top of
Main(). - Run the program—it will either start from the beginning or resume from the last saved line.
- Each time a group is finished, it updates the breakpoint file. If you stop mid-group, it will reprocess that partial group next time (since progress only saves after full groups are done).
内容的提问来源于stack exchange,提问作者Keith Hilderbrand
相关产品推荐
相关产品推荐

