You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Windows下Python多进程如何仅执行SEARCH_ENGINE函数避免重复执行脚本?

Fixing Windows Multiprocessing Spawn Behavior for Your Python Script

I totally get your frustration—Windows uses the spawn start method which re-runs your entire script in each child process, unlike Linux's fork which inherits the parent's state. Let's break down how to adjust your code so only the SEARCH_ENGINE logic runs in child processes on Windows.

The Core Issue

On Windows, every child process spawns a new Python interpreter that imports your script. Any code not wrapped in if __name__ == "__main__": will execute in every child process—including your create_file(), READ_MANY_LOG(), and select_unic_value() calls. That's why you're seeing those initialize again in children, which wastes resources and breaks your expected workflow.

Step-by-Step Fix

Here's how to restructure your code to isolate the multiprocessing logic properly:

  1. Move all initialization logic into the main guard block
    Only function definitions should live outside if __name__ == "__main__":—all function calls (like creating files, reading logs) need to run only in the parent process.
  2. Pass shared state to child processes explicitly
    Instead of relying on inherited state (which works on Linux), pass the precomputed READ_MANY_LOG results and shared return_dict as arguments to your child processes.
  3. Ensure shared objects are created in the parent
    The Manager dict needs to be initialized in the main process so children can attach to it without re-creating it.

Modified Code Example

import multiprocessing as mp

# Keep function definitions outside the main guard—these are imported but not executed in children
def create_file(output_name):
    # User-defined output file setup here
    return open(output_name, 'w')  # Return file object to parent for later use

def READ_MANY_LOG(log_paths):
    # User-defined log reading logic here
    logs_list = []
    for path in log_paths:
        with open(path, 'r') as f:
            logs_list.extend(f.readlines())
    return logs_list

def select_unic_value(logs_list):
    # User-defined unique value selection logic here
    unique_values = []
    # ... add your unique value extraction code ...
    return unique_values

def criteria(n):
    # User-defined criteria generation logic here
    return f"target_{n}"  # Example criteria—replace with your logic

def SEARCH_ENGINE(search_criteria, logs_list, return_dict):
    # User-defined search logic here
    results = []
    for line in logs_list:
        if search_criteria in line:
            results.append(line)
    return_dict[search_criteria] = results

def sort_all_logs_per_unic_value(output_file, unique_values, logs_list):
    mpc = 0
    mpa = []
    cpu_count = mp.cpu_count()
    # Create shared dict in parent process
    return_dict = mp.Manager().dict()
    
    for n in unique_values:
        search_criteria = criteria(n)
        # Pass precomputed logs and shared dict explicitly to child
        mps = mp.Process(target=SEARCH_ENGINE, args=(search_criteria, logs_list, return_dict))
        mpa.append(mps)
        mpc += 1
        
        if mpc >= cpu_count:
            # Start and wait for this batch of processes to finish
            for p in mpa:
                p.start()
            for p in mpa:
                p.join()
            
            # Write results to file now that processes are done
            for t_g in return_dict.values():
                for t_x in t_g:
                    print(t_x, file=output_file, flush=True, end='')
            print('PRINT done')
            
            # Reset for next batch
            mpc = 0
            mpa = []
            return_dict.clear()
    
    # Handle any remaining processes in the final batch
    if mpa:
        for p in mpa:
            p.start()
        for p in mpa:
            p.join()
        
        for t_g in return_dict.values():
            for t_x in t_g:
                print(t_x, file=output_file, flush=True, end='')
        print('PRINT done')

if __name__ == "__main__":
    mp.freeze_support()  # Critical for frozen executables on Windows
    
    # All initialization runs ONLY in the parent process
    output_name = "your_output_file.txt"  # Replace with user-defined name
    log_paths = ["log1.txt", "log2.txt", "log3.txt"]  # Replace with user-defined paths
    
    output_file = create_file(output_name)
    logs_list = READ_MANY_LOG(log_paths)
    unique_values = select_unic_value(logs_list)
    
    # Run the multiprocessing logic
    sort_all_logs_per_unic_value(output_file, unique_values, logs_list)
    
    # Cleanup
    output_file.close()

Key Changes Explained

  • Main Guard Block: All initialization steps (create_file, READ_MANY_LOG, select_unic_value) now run only once in the parent. Child processes import the script but skip this code entirely.
  • Explicit Argument Passing: Precomputed logs_list and shared return_dict are directly passed to SEARCH_ENGINE—children don't re-read logs or re-initialize shared state.
  • Batch Processing Fix: Fixed indentation in your original loop to ensure file writing only happens after the batch of processes has completed (your original code would have written before processes finished).
  • Cross-Platform Consistency: This structure works on both Windows and Linux, so you won't have to maintain separate versions for each OS.

内容的提问来源于stack exchange,提问作者Eugene Ti

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:03:24