You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在RSocrata中使用向量筛选多个唯一kenteken?

批量查询RSocrata数据库(基于多个kenteken值)

问题背景

你目前可以通过单条kenteken查询RSocrata数据库:

kentekens <- RSocrata::read.socrata(
  "https://opendata.rdw.nl/resource/m9d7-ebf2.json?$where=kenteken = '0001ES'")

但手里有近2000个kenteken的向量,需要实现批量查询。

实现方法

Socrata支持SQL风格的IN运算符,可直接将多个kenteken拼入查询条件。但由于2000个值数量较多,直接单次查询可能触发长度限制或API限流,分批次处理更稳妥:

1. 准备kenteken向量

先将所有目标kenteken存入一个向量:

# 替换为你实际的2000个kenteken值
kenteken_vec <- c("0001ES", "0002AB", "0003CD", ...)

2. 拆分批次

把大向量拆分为若干小批次,比如每500个一批(可根据实际情况调整):

batch_size <- 500
batches <- split(kenteken_vec, ceiling(seq_along(kenteken_vec)/batch_size))

3. 循环执行批量查询

遍历每个批次,构建查询URL并获取结果,最后合并所有批次数据:

# 初始化空列表存储各批次结果
results_list <- list()

for (batch_idx in seq_along(batches)) {
  # 将当前批次的kenteken转换为带单引号、逗号分隔的字符串
  current_kentekens <- paste0("'", batches[[batch_idx]], "'", collapse = ", ")
  # 拼接查询URL
  query_url <- paste0("https://opendata.rdw.nl/resource/m9d7-ebf2.json?$where=kenteken IN (", current_kentekens, ")")
  
  # 执行查询并保存结果
  batch_result <- RSocrata::read.socrata(query_url)
  results_list[[batch_idx]] <- batch_result
  
  # 添加1秒延迟,避免请求过于频繁触发API限流
  Sys.sleep(1)
}

# 合并所有批次结果为单个数据框
final_results <- dplyr::bind_rows(results_list)

注意事项

  • 若出现查询超时或报错,可将batch_size调小(比如200或300)
  • 延迟时间可根据API限制灵活调整,避免触发限流机制
  • 确保你的kenteken格式与数据库中完全一致(如大小写、无多余空格)

内容的提问来源于stack exchange,提问作者Raoul Van Oosten

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 03:16:07