使用multiprocessing模块递归比较列表时触发RuntimeError
Hey there! Let's break down your problem and work through solutions step by step—since you're dealing with CPU-heavy similarity comparisons on binary lists, multiprocessing should give you a nice speedup once we fix the crash issue, and we'll also cover the single-vs-multi timing comparison you need.
Before diving into fixes, let's cover the most common pitfalls that cause crashes with multiprocessing.Pool or Process:
- Missing the
if __name__ == '__main__':guard: On Windows, multiprocessing spawns new processes that re-execute your entire script. Without this guard, those child processes will accidentally launch their own child processes, creating an infinite loop that crashes your program. This is the #1 mistake people make with Python multiprocessing. - Data serialization issues:
multiprocessinguses pickle to pass data between processes. If your binary lists are stored in a non-pickleable format (or are just way too large), this can cause crashes or memory overload. - Redundant work/memory bloat: If you're passing the entire list of binary lists to every worker repeatedly, you're wasting memory by duplicating data across processes.
First, let's build a solid single-process implementation using NumPy (it's way faster than pure Python loops for binary list operations). We'll use Hamming distance (count of differing bits) as our similarity metric (similarity = total length - Hamming distance):
import numpy as np import time def generate_binary_lists(num_lists, list_len): # Generate random 0/1 lists as a NumPy array (efficient for calculations) return np.random.randint(0, 2, size=(num_lists, list_len)) def single_process_similarity(binary_arr): num_lists = binary_arr.shape[0] sim_matrix = np.zeros((num_lists, num_lists), dtype=int) # Only compute pairs once (i < j) to avoid redundant work for i in range(num_lists): for j in range(i + 1, num_lists): hamming_dist = np.sum(binary_arr[i] != binary_arr[j]) sim_score = binary_arr.shape[1] - hamming_dist sim_matrix[i][j] = sim_score sim_matrix[j][i] = sim_score return sim_matrix if __name__ == '__main__': # Test parameters NUM_LISTS = 150 LIST_LENGTH = 2000 binary_data = generate_binary_lists(NUM_LISTS, LIST_LENGTH) # Time the single-process run start = time.time() single_sim_matrix = single_process_similarity(binary_data) single_time = time.time() - start print(f"Single process time: {single_time:.2f} seconds")
Now let's rewrite this for multiprocessing, fixing the common issues. We'll use Pool to split the pair-wise calculations across workers, and avoid redundant data copying:
from multiprocessing import Pool import numpy as np import time def generate_binary_lists(num_lists, list_len): return np.random.randint(0, 2, size=(num_lists, list_len)) def compute_pair(args): # Worker function: compute similarity for a single (i,j) pair binary_arr, i, j = args hamming_dist = np.sum(binary_arr[i] != binary_arr[j]) return (i, j, binary_arr.shape[1] - hamming_dist) def multi_process_similarity(binary_arr, num_workers=None): num_lists = binary_arr.shape[0] # Generate all unique (i,j) pairs (i < j) to avoid duplicate work tasks = [(binary_arr, i, j) for i in range(num_lists) for j in range(i + 1, num_lists)] # Use a pool of workers to process tasks in parallel with Pool(num_workers) as pool: results = pool.map(compute_pair, tasks) # Build the similarity matrix from results sim_matrix = np.zeros((num_lists, num_lists), dtype=int) for i, j, sim_score in results: sim_matrix[i][j] = sim_score sim_matrix[j][i] = sim_score return sim_matrix if __name__ == '__main__': NUM_LISTS = 150 LIST_LENGTH = 2000 binary_data = generate_binary_lists(NUM_LISTS, LIST_LENGTH) # Time the multi-process run (use 4 workers, adjust based on your CPU cores) start = time.time() multi_sim_matrix = multi_process_similarity(binary_data, num_workers=4) multi_time = time.time() - start print(f"Multi-process (4 workers) time: {multi_time:.2f} seconds") # Verify results match the single-process version assert np.array_equal(single_sim_matrix, multi_sim_matrix), "Results don't match!"
You mentioned trying threading earlier—here's the deal: for CPU-heavy tasks like this, Python's Global Interpreter Lock (GIL) means threading won't give you any real speedup. The GIL only allows one thread to execute Python bytecode at a time, so threading is better for IO-heavy tasks (like network calls or file reads), not CPU-bound similarity calculations.
- Use shared memory: If your binary lists are extremely large, use
multiprocessing.Arrayor NumPy's shared memory arrays to avoid copying the entire dataset to every worker. - Bitwise operations: For shorter binary lists, you can convert each list to an integer and use XOR + bit counting to compute Hamming distance (even faster than NumPy for small lengths).
- Chunked processing: If you have tens of thousands of lists, split the task into chunks to avoid overwhelming memory with too many tasks at once.
If you still run into crashes:
- Double-check that all multiprocessing code is wrapped in
if __name__ == '__main__': - Ensure your binary data is in a pickleable format (NumPy arrays work perfectly)
- Reduce the size of your test dataset temporarily to rule out memory overload
- Try
Pool.imap_unorderedinstead ofmapto process results incrementally, reducing memory usage
内容的提问来源于stack exchange,提问作者Elliott Weaver

