如何加速API请求?Pandas数据批量API调用优化求助
Hey Riley, let's tackle this API bottleneck head-on—20k rows with 3-10s per request is brutal, and while multiprocessing helps, there are way more impactful tweaks we can make. Here's my go-to playbook for this exact scenario:
1. Prioritize Reducing API Calls (The Biggest Win)
The single most effective optimization is cutting down how many requests you send. Most APIs support batch operations, so check if yours lets you pass multiple rows of input in one call. For example:
- If your current request sends one row's data like
{"param1": "x", "param2": "y"}, see if you can send an array:[{"param1": "x", ...}, {"param1": "z", ...}] - Even if the API has a batch limit (like 50 rows per request), this drops your total requests from 20k to 400—immediately slashing total time by 90%+ in many cases.
Also, deduplicate your DataFrame first: If you have duplicate input rows, cache the API response for those duplicates and reuse them. A simple dictionary works here:
response_cache = {} def get_api_data(row): key = tuple(row.values) # Or a hashable representation of the row if key in response_cache: return response_cache[key] # Make API call here response = requests.post("https://your-api-endpoint.com", json=row.to_dict()) response_cache[key] = response.json() return response_cache[key]
2. Swap Multiprocessing for Threading/Async IO
Multiprocessing adds significant overhead for IO-bound tasks (like API calls) because each process has its own Python interpreter and memory space. Instead:
- Use
concurrent.futures.ThreadPoolExecutor: Threads are lighter, and since most of your time is waiting for API responses, threads can handle more concurrent requests without the process overhead. Example:from concurrent.futures import ThreadPoolExecutor import pandas as pd def process_row(row): # Your API call logic here (use a reusable session!) return get_api_data(row) df = pd.read_csv("your_input_data.csv") # Adjust max_workers based on API rate limits (start with 20-50) with ThreadPoolExecutor(max_workers=30) as executor: results = list(executor.map(process_row, df.itertuples(index=False))) df["api_result"] = results - Go async with
aiohttp+asyncio: For even higher concurrency, async IO lets you handle hundreds of requests at once without blocking. This is especially useful if the API allows high throughput.
3. Respect API Rate Limits (Avoid Throttling)
If you're hitting rate limits, your requests will get slowed down or blocked. Check the API docs for:
- Rate limit caps (e.g., 100 requests per minute)
- Retry-after headers, and use them to back off instead of spamming requests
- Use a library like
tenacityto add exponential backoff for retries:from tenacity import retry, stop_after_attempt, wait_exponential @retry(stop=stop_after_attempt(5), wait=wait_exponential(multiplier=1, min=2, max=10)) def get_api_data(row): response = requests.post("https://your-api-endpoint.com", json=row.to_dict()) response.raise_for_status() # Trigger retries on 4xx/5xx errors return response.json()
4. Optimize Network & Request Setup
- Reuse HTTP sessions: Use
requests.Session()oraiohttp.ClientSessionto keep TCP connections alive, avoiding the overhead of re-establishing connections for every request:session = requests.Session() def get_api_data(row): return session.post("https://your-api-endpoint.com", json=row.to_dict()).json() - Minimize payload size: Only send the parameters the API actually needs—drop any unused columns from your DataFrame before processing.
5. Cache Reusable Responses
If you run this job regularly, store API responses in a local cache (like SQLite, Redis, or even a pickle file) so you don't re-request data you've already fetched. For example, using joblib for easy caching:
from joblib import Memory memory = Memory(cachedir="./api_cache", verbose=0) @memory.cache def get_api_data(param1, param2): # Use individual params instead of rows for better cache hits response = requests.post("https://your-api-endpoint.com", json={"param1": param1, "param2": param2}) return response.json()
6. Monitor & Profile to Find Bottlenecks
Use tqdm to track progress and see how long batches take, and log slow requests to identify if certain rows are causing longer API delays. For example:
from tqdm import tqdm with ThreadPoolExecutor(max_workers=30) as executor: results = list(tqdm(executor.map(process_row, df.itertuples(index=False)), total=len(df)))
A quick note: Always test with a small subset of data first (like 100 rows) to validate any changes before scaling up to 20k.
内容的提问来源于stack exchange,提问作者Riley Hun

