如何并行拆分大文件每行单词为字符列表(避免全量加载内存)
Parallel Processing for Large Word Files Without Loading Entirely into Memory
Hey there! Let's figure out how to parallelize your word-splitting task without cramming that 3GB+ file into memory. Your current line-by-line approach is great for memory efficiency, and we can build on that with parallelism without sacrificing that benefit.
Option 1: Process File Chunks with Multiprocessing
This approach splits the file into manageable chunks, assigns each chunk to a separate process, and ensures we don't split words mid-line. Perfect for avoiding full memory loads.
Here's the code:
import multiprocessing as mp import os def process_chunk(chunk_start, chunk_size, file_path, output_queue): split_words = [] with open(file_path, 'rt') as f: # Jump to the start of the chunk f.seek(chunk_start) chunk = f.read(chunk_size) # Find the last newline to avoid cutting a word in half last_newline_idx = chunk.rfind('\n') if last_newline_idx != -1: chunk = chunk[:last_newline_idx] # Process each full line in the chunk for line in chunk.splitlines(): # Use strip() to remove unwanted newlines/whitespace; adjust if you need raw line content split_words.append(list(line.strip())) # Send results back to the main process output_queue.put(split_words) def main(): file_path = 'wordprob.txt' file_size = os.path.getsize(file_path) num_processes = mp.cpu_count() # Use all available CPU cores chunk_size = file_size // num_processes output_queue = mp.Queue() processes = [] # Spawn a process for each chunk for i in range(num_processes): start = i * chunk_size # Handle the final chunk (covers any remaining bytes) if i == num_processes - 1: chunk_size = file_size - start proc = mp.Process( target=process_chunk, args=(start, chunk_size, file_path, output_queue) ) processes.append(proc) proc.start() # Collect results from all processes final_split = [] for _ in range(num_processes): final_split.extend(output_queue.get()) # Wait for all processes to finish for proc in processes: proc.join() # Verify the first few results print("First 5 split words:", final_split[:5]) if __name__ == '__main__': main()
Why this works:
- No full memory load: Each process only reads its assigned chunk, not the entire file.
- Word safety: We find the last newline in each chunk to ensure we only process complete lines, so no words get split between chunks.
- Scalable: Uses all your CPU cores to speed up processing.
Option 2: Use concurrent.futures for Simpler Parallel Line Processing
If you prefer a more concise approach, ProcessPoolExecutor handles the parallelism boilerplate while still reading the file line-by-line (no full memory load).
Here's the code:
from concurrent.futures import ProcessPoolExecutor import os def split_word(line): # Again, adjust strip() based on whether you need to keep newlines return list(line.strip()) def main(): file_path = 'wordprob.txt' final_split = [] # Use a process pool to handle parallel processing with ProcessPoolExecutor() as executor: # File objects are iterable (return lines one at a time) with open(file_path, 'rt') as f: # executor.map processes lines in parallel while preserving order results = executor.map(split_word, f) for result in results: final_split.append(result) print("First 5 split words:", final_split[:5]) if __name__ == '__main__': main()
Why this works:
- Super clean: The
executor.mapfunction takes care of distributing lines to processes, so you don't have to manage chunks or queues manually. - Memory efficient: The main process reads lines one at a time and passes them to workers, so memory usage stays low.
- Order preserved: Results are returned in the same order as the original file, just like your original sequential code.
Key Notes to Remember
- Newline handling: Your original code uses
list(line)which includes the trailing\ncharacter. Addstrip()if you want to exclude that (adjust based on your needs). - IO bottlenecks: Parallel processing helps with CPU-bound tasks, but if your disk is slow (e.g., a mechanical HDD), you might not see huge speed gains—SSD will perform better here.
- Process count: Both examples use
cpu_count()for processes, but you can tweak this if you want to avoid overwhelming your system.
内容的提问来源于stack exchange,提问作者Keshav M
相关产品推荐
相关产品推荐

