You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何并行拆分大文件每行单词为字符列表(避免全量加载内存)

Parallel Processing for Large Word Files Without Loading Entirely into Memory

Hey there! Let's figure out how to parallelize your word-splitting task without cramming that 3GB+ file into memory. Your current line-by-line approach is great for memory efficiency, and we can build on that with parallelism without sacrificing that benefit.

Option 1: Process File Chunks with Multiprocessing

This approach splits the file into manageable chunks, assigns each chunk to a separate process, and ensures we don't split words mid-line. Perfect for avoiding full memory loads.

Here's the code:

import multiprocessing as mp
import os

def process_chunk(chunk_start, chunk_size, file_path, output_queue):
    split_words = []
    with open(file_path, 'rt') as f:
        # Jump to the start of the chunk
        f.seek(chunk_start)
        chunk = f.read(chunk_size)
        
        # Find the last newline to avoid cutting a word in half
        last_newline_idx = chunk.rfind('\n')
        if last_newline_idx != -1:
            chunk = chunk[:last_newline_idx]
        
        # Process each full line in the chunk
        for line in chunk.splitlines():
            # Use strip() to remove unwanted newlines/whitespace; adjust if you need raw line content
            split_words.append(list(line.strip()))
    
    # Send results back to the main process
    output_queue.put(split_words)

def main():
    file_path = 'wordprob.txt'
    file_size = os.path.getsize(file_path)
    num_processes = mp.cpu_count()  # Use all available CPU cores
    chunk_size = file_size // num_processes

    output_queue = mp.Queue()
    processes = []

    # Spawn a process for each chunk
    for i in range(num_processes):
        start = i * chunk_size
        # Handle the final chunk (covers any remaining bytes)
        if i == num_processes - 1:
            chunk_size = file_size - start
        
        proc = mp.Process(
            target=process_chunk,
            args=(start, chunk_size, file_path, output_queue)
        )
        processes.append(proc)
        proc.start()

    # Collect results from all processes
    final_split = []
    for _ in range(num_processes):
        final_split.extend(output_queue.get())

    # Wait for all processes to finish
    for proc in processes:
        proc.join()

    # Verify the first few results
    print("First 5 split words:", final_split[:5])

if __name__ == '__main__':
    main()

Why this works:

  • No full memory load: Each process only reads its assigned chunk, not the entire file.
  • Word safety: We find the last newline in each chunk to ensure we only process complete lines, so no words get split between chunks.
  • Scalable: Uses all your CPU cores to speed up processing.

Option 2: Use concurrent.futures for Simpler Parallel Line Processing

If you prefer a more concise approach, ProcessPoolExecutor handles the parallelism boilerplate while still reading the file line-by-line (no full memory load).

Here's the code:

from concurrent.futures import ProcessPoolExecutor
import os

def split_word(line):
    # Again, adjust strip() based on whether you need to keep newlines
    return list(line.strip())

def main():
    file_path = 'wordprob.txt'
    final_split = []

    # Use a process pool to handle parallel processing
    with ProcessPoolExecutor() as executor:
        # File objects are iterable (return lines one at a time)
        with open(file_path, 'rt') as f:
            # executor.map processes lines in parallel while preserving order
            results = executor.map(split_word, f)
            for result in results:
                final_split.append(result)

    print("First 5 split words:", final_split[:5])

if __name__ == '__main__':
    main()

Why this works:

  • Super clean: The executor.map function takes care of distributing lines to processes, so you don't have to manage chunks or queues manually.
  • Memory efficient: The main process reads lines one at a time and passes them to workers, so memory usage stays low.
  • Order preserved: Results are returned in the same order as the original file, just like your original sequential code.

Key Notes to Remember

  • Newline handling: Your original code uses list(line) which includes the trailing \n character. Add strip() if you want to exclude that (adjust based on your needs).
  • IO bottlenecks: Parallel processing helps with CPU-bound tasks, but if your disk is slow (e.g., a mechanical HDD), you might not see huge speed gains—SSD will perform better here.
  • Process count: Both examples use cpu_count() for processes, but you can tweak this if you want to avoid overwhelming your system.

内容的提问来源于stack exchange,提问作者Keshav M

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:26:37