You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C#大列表拆分:将百万级文件按99条分组断点处理

Handling Large Files with Grouping & Resume-from-Breakpoint in C#

Hey there! As someone who's been where you are (struggling with large file processing as a C# newbie), let's work through this problem step by step. We'll fix your code to handle that 10M-line password file, split it into groups of 99, pick up right where you left off if the program stops, and make it flexible enough for any list-based file.

First, Let's Break Down the Issues in Your Original Code

  • File.ReadAllLines() loads the entire 10M-line file into memory at once—this will chew up tons of RAM and might even crash your program.
  • No way to track how many lines you've already processed, so you can't resume after a break.
  • No logic to group lines into chunks of 99.

The Solution: Efficient, Resumable Group Processing

Here's a revamped version of your code that fixes all these gaps. It uses memory-friendly line-by-line reading, tracks progress with a simple breakpoint file, and lets you tweak settings for any file:

using System;
using System.IO;

namespace ListProcessingWithResume
{
    class Program
    {
        static void Main(string[] args)
        {
            // Configure these for your use case (easy to adapt to any file!)
            string inputFilePath = @"c:\tmp\10-million-password-list-top-1000000.txt";
            int groupSize = 99;
            string breakpointFilePath = @"c:\tmp\last_processed_line.txt";

            // Load progress from breakpoint file if it exists
            int linesProcessed = 0;
            if (File.Exists(breakpointFilePath))
            {
                if (int.TryParse(File.ReadAllText(breakpointFilePath), out int savedCount))
                {
                    linesProcessed = savedCount;
                    Console.WriteLine($"Resuming from line {linesProcessed + 1}...");
                }
            }

            int currentGroupCount = 0;
            // Use StreamReader to read one line at a time (memory-friendly for big files)
            using (StreamReader reader = new StreamReader(inputFilePath))
            {
                string line;
                // Skip lines we've already processed
                for (int i = 0; i < linesProcessed; i++)
                {
                    reader.ReadLine();
                }

                while ((line = reader.ReadLine()) != null)
                {
                    linesProcessed++;
                    currentGroupCount++;

                    // Add line to current group (modify this to save/process the group as needed)
                    Console.WriteLine($"Line {linesProcessed}: {line}");

                    // When we hit the group size, wrap up the group and update progress
                    if (currentGroupCount == groupSize)
                    {
                        Console.WriteLine($"--- Finished group ending at line {linesProcessed} ---");
                        
                        // Save progress so we can resume later
                        File.WriteAllText(breakpointFilePath, linesProcessed.ToString());
                        
                        currentGroupCount = 0;

                        // Optional: Add a pause if you need to throttle processing
                        // System.Threading.Thread.Sleep(100);
                    }
                }

                // Handle any leftover lines that didn't make a full group
                if (currentGroupCount > 0)
                {
                    Console.WriteLine($"--- Finished final partial group with {currentGroupCount} lines (ending at line {linesProcessed}) ---");
                    File.WriteAllText(breakpointFilePath, linesProcessed.ToString());
                }

                Console.WriteLine("All lines processed successfully!");
                // Optional: Delete the breakpoint file if you don't need it anymore
                // File.Delete(breakpointFilePath);
            }
        }
    }
}

Key Features Explained (For Newbies!)

  • Memory Efficiency: StreamReader reads one line at a time instead of loading the whole file, so even 10M lines won't bog down your system.
  • Breakpoint Resumption: The breakpointFilePath stores the last line number processed. If the program crashes or you close it, it picks up right where it left off next time.
  • Flexible Grouping: Just change the groupSize variable to split into any number of lines per group (not just 99).
  • Universal Use: Swap out inputFilePath to process any text file with one item per line—no need to rewrite the whole code.

How to Use It

  1. Set your input file path, group size, and breakpoint file path at the top of Main().
  2. Run the program—it will either start from the beginning or resume from the last saved line.
  3. Each time a group is finished, it updates the breakpoint file. If you stop mid-group, it will reprocess that partial group next time (since progress only saves after full groups are done).

内容的提问来源于stack exchange,提问作者Keith Hilderbrand

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 11:27:43