You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python入门者求教:如何将for循环转换为并行处理?

Parallelizing Your Distance & Speed Calculations in Python

Hey there! No worries at all about being early in your Python journey—we’ve all been there figuring out how to speed up repetitive calculations. Let’s break down how to turn that serial for loop into parallel processing to handle your paired latitude/longitude data more efficiently.

First, Let’s Ground the Serial Logic

Before jumping into parallelism, let’s make sure we have a solid base function for your calculations. We’ll use the Haversine formula (standard for calculating distance between two geographic points) paired with time difference to get average speed.

Here’s a basic serial implementation:

import math
import pandas as pd

def haversine(lon1, lat1, lon2, lat2):
    # Convert degrees to radians
    lon1, lat1, lon2, lat2 = map(math.radians, [lon1, lat1, lon2, lat2])
    
    # Haversine formula calculations
    dlon = lon2 - lon1
    dlat = lat2 - lat1
    a = math.sin(dlat/2)**2 + math.cos(lat1) * math.cos(lat2) * math.sin(dlon/2)**2
    c = 2 * math.asin(math.sqrt(a))
    
    # Earth radius in kilometers (use 3956 for miles)
    earth_radius = 6371
    return c * earth_radius

def calculate_speed(data_item):
    # Unpack your zipped data: ((start_lon, start_lat), (dest_lon, dest_lat), start_time, end_time)
    (start_lon, start_lat), (dest_lon, dest_lat), start_time, end_time = data_item
    
    # Calculate distance between points
    distance_km = haversine(start_lon, start_lat, dest_lon, dest_lat)
    
    # Calculate time difference in hours
    time_diff_hours = (end_time - start_time).total_seconds() / 3600
    
    # Avoid division by zero if times are identical
    if time_diff_hours == 0:
        return 0.0
    
    # Return average speed in km/h (adjust units as needed)
    return distance_km / time_diff_hours

# Example zipped data (replace with your actual data)
zipped_data = [
    ((116.397, 39.908), (116.407, 39.918), pd.Timestamp('2024-01-01 08:00'), pd.Timestamp('2024-01-01 08:30')),
    ((121.473, 31.230), (121.483, 31.240), pd.Timestamp('2024-01-01 09:00'), pd.Timestamp('2024-01-01 09:45')),
    # Add more data points here
]

# Serial for loop (your current approach)
serial_speeds = []
for item in zipped_data:
    serial_speeds.append(calculate_speed(item))

Now, Parallelize It!

Since your task is CPU-intensive (math-heavy distance calculations), we’ll use process-based parallelism (to bypass Python’s GIL, which limits thread-based parallelism for CPU-bound work). The easiest tool for beginners is concurrent.futures.ProcessPoolExecutor.

Option 1: Using concurrent.futures.ProcessPoolExecutor

This is the most beginner-friendly approach—minimal code changes from your serial loop:

from concurrent.futures import ProcessPoolExecutor

# Parallel processing with ProcessPoolExecutor
with ProcessPoolExecutor() as executor:
    # Map each item in zipped_data to the calculate_speed function
    parallel_speeds = list(executor.map(calculate_speed, zipped_data))

Key Notes for This Approach:

  • The with statement handles starting/shutting down the pool automatically (no cleanup needed!).
  • By default, it uses the number of CPU cores on your machine—you can override this with max_workers (e.g., ProcessPoolExecutor(max_workers=4)).
  • Make sure calculate_speed is a pure, pickleable function (no non-serializable objects like class instances or lambda functions inside it—this is critical for process communication).

Option 2: Using multiprocessing.Pool (Slightly Lower-Level)

If you want to use the more traditional multiprocessing module, here’s how:

from multiprocessing import Pool

# IMPORTANT: On Windows, wrap your main code in this check to avoid infinite process spawning
if __name__ == '__main__':
    with Pool() as pool:
        parallel_speeds = pool.map(calculate_speed, zipped_data)

Critical Note for Windows Users:

The if __name__ == '__main__': guard is mandatory on Windows because of how Python spawns processes. Without it, you’ll get weird errors or infinite process loops.

When to Use Parallelism?

  • Use it when you have lots of data points: Parallel processing has overhead (starting processes, moving data between them), so it won’t help with small datasets (you might even see slower performance!).
  • Stick to process pools for CPU-heavy tasks: Thread pools (ThreadPoolExecutor) are better for I/O-bound work (like downloading files), not math calculations.

Verify Your Results

Always double-check that parallel results match serial ones (accounting for tiny floating-point differences):

# Check if results are roughly equal (allow for minor float precision differences)
all(math.isclose(s, p, rel_tol=1e-9) for s, p in zip(serial_speeds, parallel_speeds))

内容的提问来源于stack exchange,提问作者alwaysaskingquestions

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:53:03