如何让Python Multiprocessing的每个进程对应文本文件中的单个邮箱
Hey there! Let's fix your multiprocessing setup so each process handles the corresponding line from your email file. The main issues in your current code are overusing global variables (which don't work across separate processes) and not properly passing the right email/index to each worker. Here are two clean solutions to achieve what you want:
Solution 1: Using multiprocessing.Process directly
This approach creates a separate process for each email line (up to the number of tasks you specify), ensuring each process gets exactly the email it needs without relying on shared globals.
import multiprocessing from colorama import Fore import time # Assume your log function is defined here def log(message): print(message) def worker(email, process_num): # This is where each process does its work with the assigned email log(f"Profile-{process_num} {Fore.GREEN} - {email}") # Add your additional task logic here (replace whatever `runit` was doing) def main(): # Read emails from file, stripping extra whitespace/newlines with open('email.txt', 'r') as f: email_list = [line.strip() for line in f if line.strip()] # Get valid user input for number of tasks while True: try: user_input = int(input(Fore.WHITE + 'How many tasks do you wanna run? [NUMBERS] \n' + Fore.RESET)) break except ValueError: print(Fore.RED + "Please enter a valid number" + Fore.RESET) # Avoid creating more processes than there are available emails num_processes = min(user_input, len(email_list)) jobs = [] for i in range(num_processes): # Pass the email and process number (i+1) directly to the worker p = multiprocessing.Process(target=worker, args=(email_list[i], i+1)) jobs.append(p) p.start() time.sleep(0.5) # Optional delay if you need staggered process startup # Wait for all processes to complete before exiting for p in jobs: p.join() if __name__ == '__main__': main()
Key improvements here:
- No globals: Each process receives its email and process number as explicit arguments, avoiding shared memory conflicts.
- Bounds checking: We cap the number of processes at the number of available emails to prevent index errors.
- Clear separation of concerns: The worker function focuses only on processing its assigned email, making the code easier to debug and modify.
Solution 2: Using multiprocessing.Pool (for scalable task management)
If you plan to handle a large number of emails, using a Pool is more efficient as it reuses processes instead of creating new ones for each task. We’ll still ensure each task maps to a specific email line:
import multiprocessing from colorama import Fore from tqdm import tqdm # Assume your log function is defined here def log(message): print(message) def pool_worker(task_data): email, task_num = task_data log(f"Profile-{task_num} {Fore.GREEN} - {email}") # Add your task logic here def main(): # Read emails from file with open('email.txt', 'r') as f: email_list = [line.strip() for line in f if line.strip()] # Get valid user input while True: try: user_input = int(input(Fore.WHITE + 'How many tasks do you wanna run? [NUMBERS] \n' + Fore.RESET)) break except ValueError: print(Fore.RED + "Please enter a valid number" + Fore.RESET) num_tasks = min(user_input, len(email_list)) # Create a list of tasks: each tuple contains the email and its corresponding task number tasks = [(email_list[i], i+1) for i in range(num_tasks)] # Use a Pool with the number of processes you want (matches num_tasks here for 1:1 mapping) with multiprocessing.Pool(processes=num_tasks) as pool: # Use tqdm to display progress bar for tasks for _ in tqdm(pool.imap(pool_worker, tasks), total=num_tasks): pass if __name__ == '__main__': main()
Notes on this approach:
- The Pool automatically manages process creation and reuse, which is more efficient for large task sets.
- We pass tuples of (email, task number) to the worker so we can still log a "profile number" that matches the email line index.
imappreserves the order of task results, so the first logged entry corresponds to the first email line (even if processing happens in parallel).
Why your original code had issues:
- Global variables: Processes have separate memory spaces, so modifying
prodor accessingemailglobally won’t work as expected across processes. - Missing task arguments: You weren’t passing the email or index to your worker functions, so they had no way to know which email to process.
- Incomplete Pool setup: In your updated code,
pool.imap_unordered(, range(...))was missing the target function, which is why it wouldn’t run.
Both solutions above will correctly assign each email line to a separate task/process, just like you wanted. Pick the one that fits your use case—direct Process creation is simpler for small numbers of tasks, while Pool is better for scalability.
内容的提问来源于stack exchange,提问作者CDNthe2nd

