You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何无需循环通过data.gov.uk API批量查询多站点?R代码优化需求

Optimizing Data Fetching from data.gov.uk Water Quality API in R

Great question—looping through each sampling point one by one is definitely going to be slow, especially with multiple sites, since each iteration fires a separate HTTP request. Let’s break down the best ways to speed this up:


1. Check for API Batch Support (Best Option)

First, verify if the data.gov.uk Water Quality API supports batch requests for multiple sampling points in a single call. Many APIs let you pass multiple IDs via a comma-separated parameter (e.g., samplingPointIds= or similar).

If this is supported, you can fetch all your data in one go, which will drastically reduce load time:

siteList <- c("TH-PTTR0016","TH-PMLE0063","TH-PTNE0072","TH-PBVR0006","TH-PTTB0002","TH-PMLR0035", "TH-PTTB0012", "TH-PTSE0072", "TH-PTNE0008", "TH-PTTB0007")

# Combine site IDs into a comma-separated string
site_ids <- paste(siteList, collapse = ",")

# Construct batch request URL (adjust parameter name if API uses something else)
batch_url <- paste0("http://environment.data.gov.uk/water-quality/measurements.csv?samplingPointIds=", site_ids, "&_limit=99999")

# Fetch all data at once
Resultall <- read.csv(batch_url)

Note: Double-check the API documentation to confirm the exact parameter name for batch sampling points—this will be the most efficient solution if available.


2. Parallel Processing (If Batch Isn’t Supported)

If the API doesn’t allow batch requests, switch from sequential looping to parallel requests to run multiple calls at the same time. Use packages like furrr (parallel-friendly purrr) or parallel to cut down waiting time:

Example with furrr:

library(furrr)
library(dplyr)

siteList <- c("TH-PTTR0016","TH-PMLE0063","TH-PTNE0072","TH-PBVR0006","TH-PTTB0002","TH-PMLR0035", "TH-PTTB0012", "TH-PTSE0072", "TH-PTNE0008", "TH-PTTB0007")

# Set up parallel session (adjust workers to match your CPU core count)
plan(multisession, workers = 4)

# Fetch data in parallel and bind results into a single data frame
Resultall <- future_map_dfr(siteList, function(site_id) {
  url <- paste0("http://environment.data.gov.uk/water-quality/id/sampling-point/", site_id, "/measurements.csv?&_limit=99999")
  read.csv(url)
})

# Reset to sequential processing when done
plan(sequential)

3. Additional Optimizations

  • Avoid repeated rbind: Your original code uses rbind in each loop, which copies the entire data frame every time and slows things down. Using bind_rows (from dplyr) or collecting results in a list first then binding once is much faster.
  • Use faster CSV readers: Replace read.csv with vroom::vroom—it’s optimized for speed with large CSV files from URLs.
  • Add a user agent: When making API requests, include a user agent string to identify your request (this helps API providers avoid flagging you as spam):
    library(httr)
    response <- GET(url, user_agent("YourProjectName/1.0 (your.email@example.com)"))
    data <- read.csv(content(response, "text"))
    
  • Handle failures gracefully: Add error handling to skip sites that fail to load instead of breaking the entire process. Use purrr::safely to wrap your fetch function.

内容的提问来源于stack exchange,提问作者IHumpheryes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:59:23