You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多进程环境下列表追加避免重复初始化的解决方案咨询

Fixing Reinitialized Global Lists in Python Multiprocessing

Hey Ryan, I totally get the frustration here—when using Python's multiprocessing, your global image_hash_list and some_dict are getting reinitialized in every child process because each process has its own isolated memory space. Let's break down why this happens and how to fix it properly.

Why Your Current Code Isn't Working

When you spawn a pool of processes, each child process gets a copy of the parent process's memory at startup. Any changes you make to image_hash_list or some_dict in a child process only affect that child's local copy—not the parent or other children. That's why it seems like the list is being reset every time.

Solution: Use multiprocessing.Manager for Shared Data Structures

The multiprocessing.Manager class creates server-side objects that can be safely shared across multiple processes. It provides thread-safe versions of common data structures like lists and dictionaries that all processes can access and modify.

Here's how to adapt your code:

Step 1: Add the if __name__ == '__main__' Guard

This is critical (especially on Windows) to prevent infinite process spawning. All your multiprocessing setup code should go inside this block.

Step 2: Create Shared Objects with Manager

Instead of using regular global lists/dicts, create shared versions via Manager(), then pass them to your worker function (using pool.starmap since we need to pass multiple arguments).

Modified Working Code

import multiprocessing as mp
import socket
import urllib.request  # Added missing import
from PIL import Image
import hashlib
import os

# Set the default timeout in seconds
timeout = 20
socket.setdefaulttimeout(timeout)

def getImages(val, image_hash_list, some_dict):
    f = open('image_files.txt', 'a')
    try:
        url = val.strip()  # Clean up leftover newlines
        local = url.split('/')[-1]
        urllib.request.urlretrieve(url, local)
        
        # Calculate image hash with proper resource handling
        with Image.open(local) as img:
            md5hash = hashlib.md5(img.tobytes())
            image_hash = md5hash.hexdigest()
        
        # Check against shared list (thread-safe operation)
        if image_hash not in image_hash_list:
            image_hash_list.append(image_hash)
            some_dict[image_hash] = 0
            f.write(f"{url}\n")
            result = 1
        else:
            print(f"Duplicate image hash detected: {image_hash}")
            result = 0
            
        os.remove(local)
        return result
    except Exception as e:
        print(f"Error processing URL {val}: {str(e)}")
        # Clean up temp file even if something goes wrong
        if os.path.exists(local):
            os.remove(local)
        return 0

if __name__ == '__main__':
    files = "Identity.txt"
    with open(files, 'r') as f:
        # Filter out empty lines and clean up whitespace
        lst = [line.strip() for line in f if line.strip()]
    
    # Create shared data structures using Manager
    with mp.Manager() as manager:
        image_hash_list = manager.list()
        some_dict = manager.dict()
        
        # Use starmap to pass multiple arguments to workers
        with mp.Pool(processes=12) as pool:
            args = [(url, image_hash_list, some_dict) for url in lst]
            res = pool.starmap(getImages, args)
    
    print(f"Processing complete! {sum(res)} new unique images were added.")

Key Changes Explained

  • Shared Objects: manager.list() and manager.dict() create proxy objects that point to a single shared server-side data structure. All processes modify the same underlying data, so no more reinitialization.
  • Argument Passing: Instead of relying on unreliable global variables, we pass shared objects directly to each worker via starmap—this is the cleanest way to share state across processes.
  • Resource Safety: Added with statements for file handling and process pools/managers to ensure resources are properly closed when done.
  • Robust Error Handling: Improved exception catching to clean up temporary files even if a download or hash calculation fails.

Alternative Approach: Centralized Processing with a Queue

If you want to avoid shared state entirely, you can have child processes send image hashes to the main process via a mp.Queue. The main process would then handle checking duplicates and updating the list. This is useful if you want to eliminate any potential race conditions (though Manager objects are already thread-safe), but the Manager approach is simpler for your use case.

Hope this fixes your issue—happy coding!

内容的提问来源于stack exchange,提问作者 Ryan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:30:38