You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言循环爬取网页报subscript out of bounds错误排查

报错原因定位

你的循环版本代码共有4处核心错误,直接触发下标越界问题:

  • 循环内每次迭代开头都执行json_data <- list(),会把上一轮循环存入的内容全部清空。第一轮n=1时给索引1赋值可正常运行,第二轮n=2时先把json_data重置为空列表,再给索引2赋值,此时索引1位置无内容,后续取值就会触发越界。
  • temp变量赋值逻辑完全错误:代码中temp<-data[n,"NA"]引用了从未定义过的data对象,属于无效笔误,和单页版本中提取url路径作为匹配key的逻辑不符。
  • R列表索引用法错误:R语言中单中括号[仅返回子列表,双中括号[[才能提取列表内存储的实际对象。你代码中json_data[n]$components、page_data[n]、info_data[n]均使用单括号索引,无法拿到嵌套的实际内容,还会触发赋值类型不匹配问题。
  • 未提前初始化存储对象:page_data、info_data均未在循环外提前定义,循环中直接对不存在的对象按索引赋值,容易出现隐式类型转换导致的异常。
修正后可运行代码

优化后去掉了不必要的全量json存储(爬500页时全存嵌套json会占用大量内存),修正了索引语法,增加了结果结构化存储:

library(rvest)
library(jsonlite)
# 目标URL列表,后续扩充到500条直接替换这个向量即可
url_list <- c("https://fortune.com/company/walmart/","https://fortune.com/company/amazon-com/")
# 提前初始化结果存储列表
res_list <- vector("list", length(url_list))

find_by_name <- function(list_data, name, elem = NULL) {
  idx <- which(sapply(list_data, \(x) x$name) == name, arr.ind = TRUE)
  stopifnot(length(idx) > 0)
  if (length(idx) > 1) { idx <- idx[1] }
  dat <- list_data[[idx]]
  if (is.null(elem)) dat else dat[[elem]]
}

for (n in seq_along(url_list)){
  # 读取当前页面预加载的JSON数据
  current_json <- read_html(url_list[n]) |>
    html_element("script#preload") |> 
    html_text() |>
    sub("\\s*window\\.__PRELOADED_STATE__ = ", "", x = _, perl = TRUE) |>
    sub(";\\s*$", "", x = _, perl = TRUE) |>
    fromJSON(simplifyVector = FALSE)
  
  # 提取当前页面对应的路径key
  current_path <- gsub(".*https://fortune.com","", url_list[n])
  current_page_data <- current_json$components$page[[current_path]]
  
  # 提取公司信息
  current_info <- current_page_data |> 
    find_by_name("body", "children") |>
    find_by_name("company-about-wrapper", "children") |>
    find_by_name("company-information", "config")
  
  # 存储需要的字段
  res_list[[n]] <- data.frame(
    url = url_list[n],
    employees = current_info$employees
  )

  # 可选:加2秒延时,避免请求过于频繁被网站封禁
  Sys.sleep(2)
}

# 合并所有结果为数据框
final_result <- do.call(rbind, res_list)
# 查看结果
print(final_result)
批量爬取建议

爬取500个页面时建议增加tryCatch容错逻辑,避免单个页面加载失败导致整个循环中断;如果出现部分页面结构不一致的情况,也可以通过容错逻辑跳过异常页面、记录失败url后续补爬。

内容的提问来源于stack exchange,提问作者Xian Zhao

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 11:09:47