如何在Python中利用更多CPU核心加速高计算脚本并恢复默认状态?
Hey Sara, totally get the frustration of waiting a week for a script to run—let’s get that 32-core server pulling its weight properly, while making sure things clean up nicely when you’re done. Here’s a step-by-step breakdown:
1. Ditch Single-Threaded Execution (Use Multiprocessing)
Python’s Global Interpreter Lock (GIL) stops threads from running CPU-bound code in parallel, so you’ll need multiprocessing instead of multithreading. The easiest way to implement this is with concurrent.futures.ProcessPoolExecutor—it handles spawning/cleaning up processes for you automatically.
Example Code:
import os from concurrent.futures import ProcessPoolExecutor def your_cpu_intensive_task(item): # Replace this with your actual task logic result = item ** 2 return result if __name__ == "__main__": # Get total available cores (leave 1-2 for system processes if needed) num_cores = os.cpu_count() # Will return 32 on your server # Or explicitly set: num_cores = 30 # Create a process pool - the 'with' block ensures clean shutdown with ProcessPoolExecutor(max_workers=num_cores) as executor: # Split your workload into chunks/items to process in parallel input_data = range(1_000_000) # Replace with your actual data # Map the task to your input data results = list(executor.map(your_cpu_intensive_task, input_data)) print("All tasks completed, resources cleaned up automatically!")
The with statement guarantees the process pool shuts down when execution finishes—no orphaned processes hanging around wasting resources.
2. Optimize Memory Usage
64GB is plenty, but poor memory management can still create bottlenecks:
- Avoid loading all data at once: If working with large datasets, process them in chunks (e.g., read a CSV line-by-line instead of loading the whole file into a DataFrame).
- Minimize inter-process data transfer: Passing large objects between processes wastes memory and time. Use shared memory (via
multiprocessing.Arrayormultiprocessing.Manager) for data that needs to be accessed by multiple processes. - Clean up unused variables: Explicitly delete variables with
deland callgc.collect()(from thegcmodule) when done with large datasets to free up memory immediately.
Example of Shared Memory:
from multiprocessing import Array, Process def process_shared_data(shared_array): for i in range(len(shared_array)): shared_array[i] *= 2 if __name__ == "__main__": # Create a shared array of integers shared_data = Array('i', [1, 2, 3, 4, 5]) p = Process(target=process_shared_data, args=(shared_data,)) p.start() p.join() # Access the modified data print(list(shared_data)) # Output: [2,4,6,8,10]
3. Optional: Adjust Process Priority (And Reset It)
If you want your script to prioritize CPU resources (without starving other system processes), you can tweak the process "nice" value. Use a try...finally block to ensure you reset it when done:
import os import psutil def set_nice_value(pid, nice_value): process = psutil.Process(pid) process.nice(nice_value) if __name__ == "__main__": original_nice = psutil.Process(os.getpid()).nice() try: # Lower nice value = higher priority (adjust based on your needs) set_nice_value(os.getpid(), -5) # Run your multiprocessing code here... finally: # Reset to original system default when finished set_nice_value(os.getpid(), original_nice)
4. Bonus: Advanced Tools for Bigger Workloads
If your script deals with massive datasets or complex computations:
- Numba: JIT-compile your CPU-intensive functions to machine code for a huge speed boost (no multiprocessing needed in some cases).
- Dask: Parallelizes pandas/numpy operations across cores, and handles out-of-core computing if your data is too big for memory.
- Cython: Convert critical parts of your code to C for near-native speed.
Final Checks
- Use
htopwhile your script runs to verify all cores are being utilized (you should see high CPU usage across most cores). - Monitor memory usage with
free -hto ensure you’re not hitting swap space (which slows things down drastically).
内容的提问来源于stack exchange,提问作者Sara

