You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中利用更多CPU核心加速高计算脚本并恢复默认状态?

Speed Up Your CPU-Intensive Python Script on a 32-Core Server

Hey Sara, totally get the frustration of waiting a week for a script to run—let’s get that 32-core server pulling its weight properly, while making sure things clean up nicely when you’re done. Here’s a step-by-step breakdown:

1. Ditch Single-Threaded Execution (Use Multiprocessing)

Python’s Global Interpreter Lock (GIL) stops threads from running CPU-bound code in parallel, so you’ll need multiprocessing instead of multithreading. The easiest way to implement this is with concurrent.futures.ProcessPoolExecutor—it handles spawning/cleaning up processes for you automatically.

Example Code:

import os
from concurrent.futures import ProcessPoolExecutor

def your_cpu_intensive_task(item):
    # Replace this with your actual task logic
    result = item ** 2
    return result

if __name__ == "__main__":
    # Get total available cores (leave 1-2 for system processes if needed)
    num_cores = os.cpu_count()  # Will return 32 on your server
    # Or explicitly set: num_cores = 30

    # Create a process pool - the 'with' block ensures clean shutdown
    with ProcessPoolExecutor(max_workers=num_cores) as executor:
        # Split your workload into chunks/items to process in parallel
        input_data = range(1_000_000)  # Replace with your actual data
        # Map the task to your input data
        results = list(executor.map(your_cpu_intensive_task, input_data))
    
    print("All tasks completed, resources cleaned up automatically!")

The with statement guarantees the process pool shuts down when execution finishes—no orphaned processes hanging around wasting resources.

2. Optimize Memory Usage

64GB is plenty, but poor memory management can still create bottlenecks:

  • Avoid loading all data at once: If working with large datasets, process them in chunks (e.g., read a CSV line-by-line instead of loading the whole file into a DataFrame).
  • Minimize inter-process data transfer: Passing large objects between processes wastes memory and time. Use shared memory (via multiprocessing.Array or multiprocessing.Manager) for data that needs to be accessed by multiple processes.
  • Clean up unused variables: Explicitly delete variables with del and call gc.collect() (from the gc module) when done with large datasets to free up memory immediately.

Example of Shared Memory:

from multiprocessing import Array, Process

def process_shared_data(shared_array):
    for i in range(len(shared_array)):
        shared_array[i] *= 2

if __name__ == "__main__":
    # Create a shared array of integers
    shared_data = Array('i', [1, 2, 3, 4, 5])
    p = Process(target=process_shared_data, args=(shared_data,))
    p.start()
    p.join()
    
    # Access the modified data
    print(list(shared_data))  # Output: [2,4,6,8,10]

3. Optional: Adjust Process Priority (And Reset It)

If you want your script to prioritize CPU resources (without starving other system processes), you can tweak the process "nice" value. Use a try...finally block to ensure you reset it when done:

import os
import psutil

def set_nice_value(pid, nice_value):
    process = psutil.Process(pid)
    process.nice(nice_value)

if __name__ == "__main__":
    original_nice = psutil.Process(os.getpid()).nice()
    try:
        # Lower nice value = higher priority (adjust based on your needs)
        set_nice_value(os.getpid(), -5)
        # Run your multiprocessing code here...
    finally:
        # Reset to original system default when finished
        set_nice_value(os.getpid(), original_nice)

4. Bonus: Advanced Tools for Bigger Workloads

If your script deals with massive datasets or complex computations:

  • Numba: JIT-compile your CPU-intensive functions to machine code for a huge speed boost (no multiprocessing needed in some cases).
  • Dask: Parallelizes pandas/numpy operations across cores, and handles out-of-core computing if your data is too big for memory.
  • Cython: Convert critical parts of your code to C for near-native speed.

Final Checks

  • Use htop while your script runs to verify all cores are being utilized (you should see high CPU usage across most cores).
  • Monitor memory usage with free -h to ensure you’re not hitting swap space (which slows things down drastically).

内容的提问来源于stack exchange,提问作者Sara

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:39:55