Python入门者求教:如何将for循环转换为并行处理?
Hey there! No worries at all about being early in your Python journey—we’ve all been there figuring out how to speed up repetitive calculations. Let’s break down how to turn that serial for loop into parallel processing to handle your paired latitude/longitude data more efficiently.
First, Let’s Ground the Serial Logic
Before jumping into parallelism, let’s make sure we have a solid base function for your calculations. We’ll use the Haversine formula (standard for calculating distance between two geographic points) paired with time difference to get average speed.
Here’s a basic serial implementation:
import math import pandas as pd def haversine(lon1, lat1, lon2, lat2): # Convert degrees to radians lon1, lat1, lon2, lat2 = map(math.radians, [lon1, lat1, lon2, lat2]) # Haversine formula calculations dlon = lon2 - lon1 dlat = lat2 - lat1 a = math.sin(dlat/2)**2 + math.cos(lat1) * math.cos(lat2) * math.sin(dlon/2)**2 c = 2 * math.asin(math.sqrt(a)) # Earth radius in kilometers (use 3956 for miles) earth_radius = 6371 return c * earth_radius def calculate_speed(data_item): # Unpack your zipped data: ((start_lon, start_lat), (dest_lon, dest_lat), start_time, end_time) (start_lon, start_lat), (dest_lon, dest_lat), start_time, end_time = data_item # Calculate distance between points distance_km = haversine(start_lon, start_lat, dest_lon, dest_lat) # Calculate time difference in hours time_diff_hours = (end_time - start_time).total_seconds() / 3600 # Avoid division by zero if times are identical if time_diff_hours == 0: return 0.0 # Return average speed in km/h (adjust units as needed) return distance_km / time_diff_hours # Example zipped data (replace with your actual data) zipped_data = [ ((116.397, 39.908), (116.407, 39.918), pd.Timestamp('2024-01-01 08:00'), pd.Timestamp('2024-01-01 08:30')), ((121.473, 31.230), (121.483, 31.240), pd.Timestamp('2024-01-01 09:00'), pd.Timestamp('2024-01-01 09:45')), # Add more data points here ] # Serial for loop (your current approach) serial_speeds = [] for item in zipped_data: serial_speeds.append(calculate_speed(item))
Now, Parallelize It!
Since your task is CPU-intensive (math-heavy distance calculations), we’ll use process-based parallelism (to bypass Python’s GIL, which limits thread-based parallelism for CPU-bound work). The easiest tool for beginners is concurrent.futures.ProcessPoolExecutor.
Option 1: Using concurrent.futures.ProcessPoolExecutor
This is the most beginner-friendly approach—minimal code changes from your serial loop:
from concurrent.futures import ProcessPoolExecutor # Parallel processing with ProcessPoolExecutor with ProcessPoolExecutor() as executor: # Map each item in zipped_data to the calculate_speed function parallel_speeds = list(executor.map(calculate_speed, zipped_data))
Key Notes for This Approach:
- The
withstatement handles starting/shutting down the pool automatically (no cleanup needed!). - By default, it uses the number of CPU cores on your machine—you can override this with
max_workers(e.g.,ProcessPoolExecutor(max_workers=4)). - Make sure
calculate_speedis a pure, pickleable function (no non-serializable objects like class instances or lambda functions inside it—this is critical for process communication).
Option 2: Using multiprocessing.Pool (Slightly Lower-Level)
If you want to use the more traditional multiprocessing module, here’s how:
from multiprocessing import Pool # IMPORTANT: On Windows, wrap your main code in this check to avoid infinite process spawning if __name__ == '__main__': with Pool() as pool: parallel_speeds = pool.map(calculate_speed, zipped_data)
Critical Note for Windows Users:
The if __name__ == '__main__': guard is mandatory on Windows because of how Python spawns processes. Without it, you’ll get weird errors or infinite process loops.
When to Use Parallelism?
- Use it when you have lots of data points: Parallel processing has overhead (starting processes, moving data between them), so it won’t help with small datasets (you might even see slower performance!).
- Stick to process pools for CPU-heavy tasks: Thread pools (
ThreadPoolExecutor) are better for I/O-bound work (like downloading files), not math calculations.
Verify Your Results
Always double-check that parallel results match serial ones (accounting for tiny floating-point differences):
# Check if results are roughly equal (allow for minor float precision differences) all(math.isclose(s, p, rel_tol=1e-9) for s, p in zip(serial_speeds, parallel_speeds))
内容的提问来源于stack exchange,提问作者alwaysaskingquestions

