You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Bash中批量并行执行Python脚本的优化方案咨询

Optimized Ways to Process Hundreds of Batch Files with Controlled Parallelism in Python

Great question! Handling large batches of batch files with a fixed number of concurrent tasks is super common, and Python has some really solid, low-fuss tools to make this smooth. Here are my top recommended approaches, ordered by simplicity and practicality:

1. Use concurrent.futures.ProcessPoolExecutor (Best for Most Cases)

This is my go-to choice—it’s part of Python’s standard library, requires no extra installs, and handles parallelism with minimal code. Perfect when you just need to cap concurrent tasks at 10.

Here’s a complete example:

import os
from concurrent.futures import ProcessPoolExecutor, as_completed

def process_batch_file(file_path):
    """Your existing batch file processing logic here"""
    try:
        # Replace this with your actual processing code (e.g., subprocess.run)
        print(f"Processing {os.path.basename(file_path)}...")
        # Example: Run the batch file
        # result = subprocess.run(file_path, check=True, capture_output=True, text=True)
        # return (file_path, "SUCCESS", result.stdout)
        return (file_path, "SUCCESS", "Processing completed")
    except Exception as e:
        return (file_path, "FAILED", str(e))

def main():
    batches_dir = "./batches"
    # Get all .bat/.cmd files in the directory
    batch_files = [
        os.path.join(batches_dir, f)
        for f in os.listdir(batches_dir)
        if f.lower().endswith((".bat", ".cmd"))
    ]

    # Process with max 10 concurrent tasks
    with ProcessPoolExecutor(max_workers=10) as executor:
        # Submit all tasks
        futures = {executor.submit(process_batch_file, file): file for file in batch_files}
        
        # Track results as they complete
        for future in as_completed(futures):
            file_path, status, message = future.result()
            print(f"[{status}] {os.path.basename(file_path)}: {message}")

if __name__ == "__main__":
    main()
  • Why this works: The executor manages the worker pool automatically, distributing tasks to idle workers and handling cleanup.
  • Pro tip: If your processing is IO-bound (e.g., reading/writing files, waiting on external commands), swap ProcessPoolExecutor with ThreadPoolExecutor—it has lower overhead since it uses threads instead of separate processes.

2. Use multiprocessing.Pool (More Low-Level Control)

If you prefer a slightly more explicit approach, Python’s multiprocessing module’s Pool class is another standard library option. It’s great if you need to tweak pool initialization or use different task submission methods.

Example code:

import os
import multiprocessing

def process_batch_file(file_path):
    # Same processing function as above
    try:
        print(f"Processing {os.path.basename(file_path)}...")
        return (file_path, "SUCCESS", "Processing completed")
    except Exception as e:
        return (file_path, "FAILED", str(e))

def main():
    batches_dir = "./batches"
    batch_files = [
        os.path.join(batches_dir, f)
        for f in os.listdir(batches_dir)
        if f.lower().endswith((".bat", ".cmd"))
    ]

    # Initialize pool with 10 workers
    with multiprocessing.Pool(processes=10) as pool:
        # Use imap_unordered to get results as they finish (instead of waiting for all)
        for result in pool.imap_unordered(process_batch_file, batch_files):
            file_path, status, message = result
            print(f"[{status}] {os.path.basename(file_path)}: {message}")

if __name__ == "__main__":
    main()

3. Custom Task Queue (For Advanced Control)

If you need things like task prioritization, retry logic, or real-time status tracking, you can build a custom queue using multiprocessing.Queue and worker processes. This is overkill for basic cases, but useful for complex workflows.

Key Optimizations to Add Regardless of Approach

  • Filter valid files upfront: Make sure you’re only processing actual batch files (avoid directories, temporary files, etc.) like the examples above.
  • Log everything: Replace print statements with a proper logging setup (using Python’s logging module) so you can debug failures later.
  • Handle edge cases: Add checks for file existence, read permissions, and empty files before processing.
  • Throttle if needed: If your batch files stress system resources (CPU, disk, network), you could add small delays or adjust the max worker count dynamically, but 10 is usually safe for most systems.

Pick the approach that fits your workflow—ProcessPoolExecutor is almost always the best starting point because it’s simple, robust, and requires zero extra dependencies.

内容的提问来源于stack exchange,提问作者MrD

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:44:58