You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3嵌套循环优化:NHS坐标匹配任务并行方案咨询

Optimizing NHS Distance Calculation: From 3-4 Days to Hours

Hey there, let's tackle this distance calculation bottleneck head-on—3-4 days of runtime is way too slow, and we can get this down to a matter of hours with the right optimizations. Let's break down the best approach and walk through a working code example.

First: Why Your Current Code Is Slow

Your double loop is doing an O(n*m) calculation (8,000 clinics × 181,000 regions = ~1.45 billion distance checks) which is inherently heavy, but the main bottlenecks are:

  • Single-threaded execution: You're only using one of your 6 CPU cores.
  • Geopy's overhead: geopy.distance is flexible but adds extra layers that slow down raw distance math.
  • Pandas index lookups in loops: Repeatedly accessing df3["Lat"][i] inside loops adds unnecessary overhead.

Best Optimization Approach: Multiprocessing with a Faster Distance Library

Here's why this beats your other options:

  • itertools: Doesn't change the time complexity—it just rearranges how you write the loop, so no speed gain.
  • concurrent.futures.ThreadPoolExecutor: Useless for CPU-heavy tasks like distance calculation, thanks to Python's Global Interpreter Lock (GIL) which prevents true parallelism in threads.
  • multiprocessing.Pool (or concurrent.futures.ProcessPoolExecutor): Creates separate Python processes, each with its own GIL, so you can leverage all 6 of your CPU cores.

Expected Performance Gain

With 6 cores, you'll get roughly 4-5x speedup (accounting for process communication overhead). Pair that with a faster distance library like haversine (instead of geopy), and you can cut runtime from 3-4 days to 4-6 hours—a massive improvement.

Step-by-Step Optimized Code

First, install the haversine library if you haven't already:

pip install haversine

Then, here's the revised code:

1. Preprocess Data for Speed

We'll extract the necessary data from pandas into lists/arrays to avoid slow index lookups in loops:

import numpy as np
from haversine import haversine, Unit
from multiprocessing import Pool

# Preprocess clinic data: extract Organisation Code + coordinates as floats
clinic_data = list(zip(
    df3["Organisation_Code"],
    df3["Lat"].astype(float),
    df3["Lon"].astype(float)
))

# Preprocess OA data: extract code + coordinates as floats
oa_data = list(zip(
    oa_df["code"],
    oa_df["lat"].astype(float),
    oa_df["lon"].astype(float)
))

2. Define a Parallelizable Clinic Processing Function

This function handles all distance checks for a single clinic, returning only the results we care about (and any failed records):

def process_single_clinic(clinic_info):
    org_code, clinic_lat, clinic_lon = clinic_info
    clinic_coords = (clinic_lat, clinic_lon)
    matched_records = []
    failed_records = []

    for oa_code, oa_lat, oa_lon in oa_data:
        try:
            # Use haversine for faster distance calculation (WGS84 coordinates)
            distance_km = haversine(clinic_coords, (oa_lat, oa_lon), unit=Unit.KILOMETERS)
            if distance_km <= 3:
                matched_records.append([org_code, oa_code, distance_km])
        except Exception as e:
            failed_records.append([
                org_code, oa_code, clinic_lat, clinic_lon, oa_lat, oa_lon
            ])
            print(f"Failed match for {org_code} & {oa_code}: {str(e)}")
    
    return matched_records, failed_records

3. Run with Multiprocessing Pool

We'll use a pool of processes to handle clinics in parallel, then aggregate results in the main process:

if __name__ == "__main__":
    surgery_oa_list = []
    failed = []

    # Use 5 processes (leave 1 core free for system tasks)
    with Pool(processes=5) as pool:
        # Iterate through results as they're completed
        for clinic_matches, clinic_fails in pool.imap(process_single_clinic, clinic_data):
            surgery_oa_list.extend(clinic_matches)
            failed.extend(clinic_fails)
            print(f"Processed a clinic | Added {len(clinic_matches)} new hits")
    
    print("\n=== Final Results ===")
    print(f"Total matched records: {len(surgery_oa_list)}")
    print(f"Total failed records: {len(failed)}")
    print("Done!")

Extra Tips for Even More Speed

  • Batch processing with numpy: If you're comfortable with numpy, you can vectorize distance calculations for each clinic (compute all OA distances at once with array operations), which can squeeze out another 2-3x speedup.
  • Filter invalid coordinates first: Pre-check for NaN/non-numeric coordinates in df3 and oa_df before processing, so you don't waste time handling exceptions in loops.
  • Use pool.imap_unordered instead of imap: If you don't care about processing order (which you said you don't), imap_unordered can return results faster as they finish, instead of waiting for clinics in sequence.

内容的提问来源于stack exchange,提问作者Jonathan Francis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 13:17:43