You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何识别待抓取表格类型?ArcGIS表格能否用rvest抓取?

问题分析与解决方案

你的当前思路不正确,核心原因如下:

  • ArcGIS Online生成的页面表格是动态渲染的前端组件,并非传统HTML <table> 元素,页面静态源码里没有存储表格的实际数据,所以SelectorGadget找不到有效选择器,rvest的html_table()也无法提取到内容。
  • 26万行属于大数据量,即便能抓到前端渲染的内容,也会因分页加载、前端性能限制等问题无法完整获取。

正确思路:直接调用ArcGIS REST API获取数据

这个页面展示的是ArcGIS要素服务(Feature Service),绕过前端页面、直接调用官方REST API接口获取原始数据,才是最可靠、高效的方式。

步骤1:获取要素服务的REST端点URL

  1. 打开你提供的页面,切换到「Data」标签页
  2. 找到对应图层(比如该item关联的Groundwater_Levels图层),通过「View」选项或页面源码搜索FeatureServer关键词,拿到格式如下的REST服务URL:
    https://gispublic.waterboards.ca.gov/arcgis/rest/services/Public/Groundwater_Levels/FeatureServer/0
    

步骤2:用R批量获取数据

ArcGIS REST API默认每次最多返回1000条记录,需要通过offset参数分页请求,以下是两种实用实现方式:

方式1:用jsonlite+httr手动处理分页

library(jsonlite)
library(httr)

# 替换为你找到的要素服务URL
service_url <- "https://gispublic.waterboards.ca.gov/arcgis/rest/services/Public/Groundwater_Levels/FeatureServer/0/query"

# 先获取总记录数
total_count <- GET(service_url, query = list(
  where = "1=1",
  returnCountOnly = "true",
  f = "json"
)) %>% content("text") %>% fromJSON() %>% .$count

# 分页请求所有数据
batch_size <- 1000
all_data <- data.frame()

for (offset in seq(0, total_count - 1, batch_size)) {
  response <- GET(service_url, query = list(
    where = "1=1",
    outFields = "*",
    returnGeometry = "false", # 不需要空间数据时设为false,加快请求速度
    resultOffset = offset,
    resultRecordCount = batch_size,
    f = "json"
  ))
  
  batch_data <- content(response, "text") %>% fromJSON() %>% .$features %>% .$attributes
  all_data <- rbind(all_data, batch_data)
  
  # 打印进度
  cat(sprintf("已获取 %d/%d 条记录\n", nrow(all_data), total_count))
}

# 保存结果
write.csv(all_data, "groundwater_levels.csv", row.names = FALSE)

方式2:用arcgisbinding包(更简便)

arcgisbinding是R与ArcGIS交互的专用包,能自动处理分页逻辑,适合大数据量场景:

library(arcgisbinding)

# 初始化ArcGIS连接
arc.check_product()

# 替换为要素服务URL
service_url <- "https://gispublic.waterboards.ca.gov/arcgis/rest/services/Public/Groundwater_Levels/FeatureServer/0"

# 直接读取全量数据(自动分页)
all_data <- arc.open(service_url) %>% arc.select()

# 保存结果
write.csv(all_data, "groundwater_levels.csv", row.names = FALSE)

关键提示

  • API方式从数据源直接拉取数据,完全规避前端渲染的不稳定问题
  • 如果需要空间地理数据,将returnGeometry设为true即可,arcgisbinding会自动处理空间字段
  • 处理26万行数据时,建议分批次保存或使用data.table优化内存占用

内容的提问来源于stack exchange,提问作者dbo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 16:17:30