多进程环境下列表追加避免重复初始化的解决方案咨询
Hey Ryan, I totally get the frustration here—when using Python's multiprocessing, your global image_hash_list and some_dict are getting reinitialized in every child process because each process has its own isolated memory space. Let's break down why this happens and how to fix it properly.
Why Your Current Code Isn't Working
When you spawn a pool of processes, each child process gets a copy of the parent process's memory at startup. Any changes you make to image_hash_list or some_dict in a child process only affect that child's local copy—not the parent or other children. That's why it seems like the list is being reset every time.
Solution: Use multiprocessing.Manager for Shared Data Structures
The multiprocessing.Manager class creates server-side objects that can be safely shared across multiple processes. It provides thread-safe versions of common data structures like lists and dictionaries that all processes can access and modify.
Here's how to adapt your code:
Step 1: Add the if __name__ == '__main__' Guard
This is critical (especially on Windows) to prevent infinite process spawning. All your multiprocessing setup code should go inside this block.
Step 2: Create Shared Objects with Manager
Instead of using regular global lists/dicts, create shared versions via Manager(), then pass them to your worker function (using pool.starmap since we need to pass multiple arguments).
Modified Working Code
import multiprocessing as mp import socket import urllib.request # Added missing import from PIL import Image import hashlib import os # Set the default timeout in seconds timeout = 20 socket.setdefaulttimeout(timeout) def getImages(val, image_hash_list, some_dict): f = open('image_files.txt', 'a') try: url = val.strip() # Clean up leftover newlines local = url.split('/')[-1] urllib.request.urlretrieve(url, local) # Calculate image hash with proper resource handling with Image.open(local) as img: md5hash = hashlib.md5(img.tobytes()) image_hash = md5hash.hexdigest() # Check against shared list (thread-safe operation) if image_hash not in image_hash_list: image_hash_list.append(image_hash) some_dict[image_hash] = 0 f.write(f"{url}\n") result = 1 else: print(f"Duplicate image hash detected: {image_hash}") result = 0 os.remove(local) return result except Exception as e: print(f"Error processing URL {val}: {str(e)}") # Clean up temp file even if something goes wrong if os.path.exists(local): os.remove(local) return 0 if __name__ == '__main__': files = "Identity.txt" with open(files, 'r') as f: # Filter out empty lines and clean up whitespace lst = [line.strip() for line in f if line.strip()] # Create shared data structures using Manager with mp.Manager() as manager: image_hash_list = manager.list() some_dict = manager.dict() # Use starmap to pass multiple arguments to workers with mp.Pool(processes=12) as pool: args = [(url, image_hash_list, some_dict) for url in lst] res = pool.starmap(getImages, args) print(f"Processing complete! {sum(res)} new unique images were added.")
Key Changes Explained
- Shared Objects:
manager.list()andmanager.dict()create proxy objects that point to a single shared server-side data structure. All processes modify the same underlying data, so no more reinitialization. - Argument Passing: Instead of relying on unreliable global variables, we pass shared objects directly to each worker via
starmap—this is the cleanest way to share state across processes. - Resource Safety: Added
withstatements for file handling and process pools/managers to ensure resources are properly closed when done. - Robust Error Handling: Improved exception catching to clean up temporary files even if a download or hash calculation fails.
Alternative Approach: Centralized Processing with a Queue
If you want to avoid shared state entirely, you can have child processes send image hashes to the main process via a mp.Queue. The main process would then handle checking duplicates and updating the list. This is useful if you want to eliminate any potential race conditions (though Manager objects are already thread-safe), but the Manager approach is simpler for your use case.
Hope this fixes your issue—happy coding!
内容的提问来源于stack exchange,提问作者 Ryan

