如何对JSON API循环查询进行多进程优化以缩短加载时间?
Absolutely! Multi-processing (or even multi-threading, since this is I/O-bound work) is a perfect fix here—your current loop waits on each API response one at a time, wasting tons of idle time while your computer sits around waiting for network calls to finish. Let's walk through how to optimize this.
Why your current code is slow
Right now, your code sends one API request, waits for it to fully complete, then moves to the next. Since API calls spend most of their time waiting for a server response (not using your CPU), you're leaving massive performance on the table by not running these requests in parallel.
Multi-processing implementation with concurrent.futures
The easiest way to parallelize this in Python is using concurrent.futures.ProcessPoolExecutor (or ThreadPoolExecutor for lighter overhead, since this is I/O-bound work). Here's how to refactor your code:
First, extract the API fetch logic into a standalone function—this makes it easy to run in parallel:
import requests from concurrent.futures import ProcessPoolExecutor def fetch_single_train_data(train_id, target_date): auth_creds = ('username', 'password') api_url = f'https://api.com/{train_id}/{target_date}' try: response = requests.get(api_url, auth=auth_creds) response.raise_for_status() # Catch HTTP errors like 404/500 train_service_json = response.json() # Replicate your original station list logic here list_of_all_stations = [] # Add your code to populate list_of_all_stations from train_service_json # Return successful result + train ID to map back to your original list return (train_id, list_of_all_stations, None) except Exception as e: # Handle all possible errors: network issues, bad JSON, missing keys, etc. error_msg = str(e) # Check for the specific 'error' key from your original code if 'train_service_json' in locals() and 'error' in train_service_json: error_msg = str(train_service_json['error']) return (train_id, None, error_msg)
Then, run all requests in parallel in your main logic:
# Assuming your pre-defined variables exist: allTrainDataList, mergeAllTrainDataList, date processed_results = [] # Adjust max_workers based on API rate limits (don't flood the API!) # A good starting point is 4-8, or check the API docs for allowed concurrency max_parallel_workers = 6 with ProcessPoolExecutor(max_workers=max_parallel_workers) as executor: # Submit all API fetch tasks to the pool future_tasks = [ executor.submit(fetch_single_train_data, train_id, date) for train_id in mergeAllTrainDataList ] # Collect results as they finish for task in future_tasks: train_id, stations, error = task.result() if error: print(f"Warning: Failed to fetch data for train {train_id} - {error}") else: # Map the result back to your original allTrainDataList # Find the index of the train ID in mergeAllTrainDataList train_index = mergeAllTrainDataList.index(train_id) processed_results.append((train_index, stations)) # Update your original list with the fetched data for idx, stations in processed_results: # Adjust this line to match your actual data structure in allTrainDataList allTrainDataList[idx]['stations'] = stations
Key notes to keep in mind
- Threading vs Processing: Since this is I/O-bound work (waiting on network),
ThreadPoolExecutorwill work just as well and has less overhead (no process memory duplication). Feel free to swapProcessPoolExecutorwithThreadPoolExecutorto test which is faster for your use case. - API Rate Limits: Don't set
max_workerstoo high! Most APIs have rate limits or concurrency caps—check the API documentation to avoid getting blocked. - Robust Error Handling: The refactored code catches more edge cases (network timeouts, HTTP errors, invalid JSON) than your original try/except, which only checked for an 'error' key.
- Result Mapping: We return the
train_idwith each result so you can correctly map the fetched data back to your originalallTrainDataList—this is important because parallel tasks don't finish in order.
内容的提问来源于stack exchange,提问作者Gray Ridley

