如何无需循环通过data.gov.uk API批量查询多站点?R代码优化需求
Great question—looping through each sampling point one by one is definitely going to be slow, especially with multiple sites, since each iteration fires a separate HTTP request. Let’s break down the best ways to speed this up:
1. Check for API Batch Support (Best Option)
First, verify if the data.gov.uk Water Quality API supports batch requests for multiple sampling points in a single call. Many APIs let you pass multiple IDs via a comma-separated parameter (e.g., samplingPointIds= or similar).
If this is supported, you can fetch all your data in one go, which will drastically reduce load time:
siteList <- c("TH-PTTR0016","TH-PMLE0063","TH-PTNE0072","TH-PBVR0006","TH-PTTB0002","TH-PMLR0035", "TH-PTTB0012", "TH-PTSE0072", "TH-PTNE0008", "TH-PTTB0007") # Combine site IDs into a comma-separated string site_ids <- paste(siteList, collapse = ",") # Construct batch request URL (adjust parameter name if API uses something else) batch_url <- paste0("http://environment.data.gov.uk/water-quality/measurements.csv?samplingPointIds=", site_ids, "&_limit=99999") # Fetch all data at once Resultall <- read.csv(batch_url)
Note: Double-check the API documentation to confirm the exact parameter name for batch sampling points—this will be the most efficient solution if available.
2. Parallel Processing (If Batch Isn’t Supported)
If the API doesn’t allow batch requests, switch from sequential looping to parallel requests to run multiple calls at the same time. Use packages like furrr (parallel-friendly purrr) or parallel to cut down waiting time:
Example with furrr:
library(furrr) library(dplyr) siteList <- c("TH-PTTR0016","TH-PMLE0063","TH-PTNE0072","TH-PBVR0006","TH-PTTB0002","TH-PMLR0035", "TH-PTTB0012", "TH-PTSE0072", "TH-PTNE0008", "TH-PTTB0007") # Set up parallel session (adjust workers to match your CPU core count) plan(multisession, workers = 4) # Fetch data in parallel and bind results into a single data frame Resultall <- future_map_dfr(siteList, function(site_id) { url <- paste0("http://environment.data.gov.uk/water-quality/id/sampling-point/", site_id, "/measurements.csv?&_limit=99999") read.csv(url) }) # Reset to sequential processing when done plan(sequential)
3. Additional Optimizations
- Avoid repeated
rbind: Your original code usesrbindin each loop, which copies the entire data frame every time and slows things down. Usingbind_rows(fromdplyr) or collecting results in a list first then binding once is much faster. - Use faster CSV readers: Replace
read.csvwithvroom::vroom—it’s optimized for speed with large CSV files from URLs. - Add a user agent: When making API requests, include a user agent string to identify your request (this helps API providers avoid flagging you as spam):
library(httr) response <- GET(url, user_agent("YourProjectName/1.0 (your.email@example.com)")) data <- read.csv(content(response, "text")) - Handle failures gracefully: Add error handling to skip sites that fail to load instead of breaking the entire process. Use
purrr::safelyto wrap your fetch function.
内容的提问来源于stack exchange,提问作者IHumpheryes

