R语言循环优化:基于API限制实现批量300条提取后延时执行
Optimizing R API Calls for Batch Processing & Rate Limiting
Got it, let's fix this to make the most of your API's rate limit! Instead of adding a 1.5s delay after every single call, we'll batch your requests into groups of 300 (matching the API's 5-minute limit), process each batch as fast as possible, then wait 5 minutes before the next batch. Here's a clean, efficient implementation:
Step-by-Step Refactored Code
# Your patient ID vector (example provided) patient_number <- c("10523563", "08989388", "12338490", "11362265", "08732885", "06464116", "02841570", "09966472", "11260599", "06797991", "11047641", "09538121", "01298039", "00875727", "SC464438", "11361061", "SC294938", "02450551", "09447238", "12229772", "10562525", "11834746", "08226516", "11480020", "SC255490", "11794931", "01298989", "09782645") # Split IDs into batches of 300 (last batch will be smaller if total isn't a multiple) patient_batches <- split(patient_number, ceiling(seq_along(patient_number)/300)) # Initialize empty list to store all patient records patients <- list() # Process each batch for (batch_idx in seq_along(patient_batches)) { current_batch <- patient_batches[[batch_idx]] # Process all IDs in the current batch batch_records <- lapply(current_batch, function(pid) { patient_ID_extract(pid, api_key) # Optional: Add a tiny delay here (e.g., Sys.sleep(0.1)) if API rejects rapid calls }) # Merge batch results into the main list patients <- c(patients, batch_records) # Wait 5 minutes only if there are more batches to process if (batch_idx != length(patient_batches)) { message(paste("Completed batch", batch_idx, "- waiting 5 minutes for API limit reset...")) Sys.sleep(5 * 60) # Convert 5 minutes to seconds } }
Key Improvements Explained
- Batch Grouping: The
split()function withceiling(seq_along(...)/300)neatly divides your patient IDs into chunks of exactly 300, which aligns perfectly with the API's rate limit rules. - No Unnecessary Delays: We removed the per-request 1.5s sleep, so each batch processes as fast as the API can handle it—this cuts down total runtime significantly.
- Smart Waiting: We only pause after batches that aren't the last one, so you don't waste time waiting when there's no more data to fetch.
- Optional Per-Request Delay: If you notice the API is rejecting rapid consecutive calls, you can add a tiny
Sys.sleep(0.1)inside thelapplyfunction to space out requests within a batch without killing efficiency.
This approach ensures you're making full use of the API's allowed calls each 5-minute window, while respecting the rate limit to avoid getting blocked.
内容的提问来源于stack exchange,提问作者codedancer
相关产品推荐
相关产品推荐

