Python3嵌套循环优化:NHS坐标匹配任务并行方案咨询
Hey there, let's tackle this distance calculation bottleneck head-on—3-4 days of runtime is way too slow, and we can get this down to a matter of hours with the right optimizations. Let's break down the best approach and walk through a working code example.
First: Why Your Current Code Is Slow
Your double loop is doing an O(n*m) calculation (8,000 clinics × 181,000 regions = ~1.45 billion distance checks) which is inherently heavy, but the main bottlenecks are:
- Single-threaded execution: You're only using one of your 6 CPU cores.
- Geopy's overhead:
geopy.distanceis flexible but adds extra layers that slow down raw distance math. - Pandas index lookups in loops: Repeatedly accessing
df3["Lat"][i]inside loops adds unnecessary overhead.
Best Optimization Approach: Multiprocessing with a Faster Distance Library
Here's why this beats your other options:
itertools: Doesn't change the time complexity—it just rearranges how you write the loop, so no speed gain.concurrent.futures.ThreadPoolExecutor: Useless for CPU-heavy tasks like distance calculation, thanks to Python's Global Interpreter Lock (GIL) which prevents true parallelism in threads.multiprocessing.Pool(orconcurrent.futures.ProcessPoolExecutor): Creates separate Python processes, each with its own GIL, so you can leverage all 6 of your CPU cores.
Expected Performance Gain
With 6 cores, you'll get roughly 4-5x speedup (accounting for process communication overhead). Pair that with a faster distance library like haversine (instead of geopy), and you can cut runtime from 3-4 days to 4-6 hours—a massive improvement.
Step-by-Step Optimized Code
First, install the haversine library if you haven't already:
pip install haversine
Then, here's the revised code:
1. Preprocess Data for Speed
We'll extract the necessary data from pandas into lists/arrays to avoid slow index lookups in loops:
import numpy as np from haversine import haversine, Unit from multiprocessing import Pool # Preprocess clinic data: extract Organisation Code + coordinates as floats clinic_data = list(zip( df3["Organisation_Code"], df3["Lat"].astype(float), df3["Lon"].astype(float) )) # Preprocess OA data: extract code + coordinates as floats oa_data = list(zip( oa_df["code"], oa_df["lat"].astype(float), oa_df["lon"].astype(float) ))
2. Define a Parallelizable Clinic Processing Function
This function handles all distance checks for a single clinic, returning only the results we care about (and any failed records):
def process_single_clinic(clinic_info): org_code, clinic_lat, clinic_lon = clinic_info clinic_coords = (clinic_lat, clinic_lon) matched_records = [] failed_records = [] for oa_code, oa_lat, oa_lon in oa_data: try: # Use haversine for faster distance calculation (WGS84 coordinates) distance_km = haversine(clinic_coords, (oa_lat, oa_lon), unit=Unit.KILOMETERS) if distance_km <= 3: matched_records.append([org_code, oa_code, distance_km]) except Exception as e: failed_records.append([ org_code, oa_code, clinic_lat, clinic_lon, oa_lat, oa_lon ]) print(f"Failed match for {org_code} & {oa_code}: {str(e)}") return matched_records, failed_records
3. Run with Multiprocessing Pool
We'll use a pool of processes to handle clinics in parallel, then aggregate results in the main process:
if __name__ == "__main__": surgery_oa_list = [] failed = [] # Use 5 processes (leave 1 core free for system tasks) with Pool(processes=5) as pool: # Iterate through results as they're completed for clinic_matches, clinic_fails in pool.imap(process_single_clinic, clinic_data): surgery_oa_list.extend(clinic_matches) failed.extend(clinic_fails) print(f"Processed a clinic | Added {len(clinic_matches)} new hits") print("\n=== Final Results ===") print(f"Total matched records: {len(surgery_oa_list)}") print(f"Total failed records: {len(failed)}") print("Done!")
Extra Tips for Even More Speed
- Batch processing with numpy: If you're comfortable with numpy, you can vectorize distance calculations for each clinic (compute all OA distances at once with array operations), which can squeeze out another 2-3x speedup.
- Filter invalid coordinates first: Pre-check for NaN/non-numeric coordinates in
df3andoa_dfbefore processing, so you don't waste time handling exceptions in loops. - Use
pool.imap_unorderedinstead ofimap: If you don't care about processing order (which you said you don't),imap_unorderedcan return results faster as they finish, instead of waiting for clinics in sequence.
内容的提问来源于stack exchange,提问作者Jonathan Francis

