R语言循环爬取网页报subscript out of bounds错误排查
报错原因定位
你的循环版本代码共有4处核心错误,直接触发下标越界问题:
- 循环内每次迭代开头都执行
json_data <- list(),会把上一轮循环存入的内容全部清空。第一轮n=1时给索引1赋值可正常运行,第二轮n=2时先把json_data重置为空列表,再给索引2赋值,此时索引1位置无内容,后续取值就会触发越界。 temp变量赋值逻辑完全错误:代码中temp<-data[n,"NA"]引用了从未定义过的data对象,属于无效笔误,和单页版本中提取url路径作为匹配key的逻辑不符。- R列表索引用法错误:R语言中单中括号
[仅返回子列表,双中括号[[才能提取列表内存储的实际对象。你代码中json_data[n]$components、page_data[n]、info_data[n]均使用单括号索引,无法拿到嵌套的实际内容,还会触发赋值类型不匹配问题。 - 未提前初始化存储对象:
page_data、info_data均未在循环外提前定义,循环中直接对不存在的对象按索引赋值,容易出现隐式类型转换导致的异常。
修正后可运行代码
优化后去掉了不必要的全量json存储(爬500页时全存嵌套json会占用大量内存),修正了索引语法,增加了结果结构化存储:
library(rvest) library(jsonlite) # 目标URL列表,后续扩充到500条直接替换这个向量即可 url_list <- c("https://fortune.com/company/walmart/","https://fortune.com/company/amazon-com/") # 提前初始化结果存储列表 res_list <- vector("list", length(url_list)) find_by_name <- function(list_data, name, elem = NULL) { idx <- which(sapply(list_data, \(x) x$name) == name, arr.ind = TRUE) stopifnot(length(idx) > 0) if (length(idx) > 1) { idx <- idx[1] } dat <- list_data[[idx]] if (is.null(elem)) dat else dat[[elem]] } for (n in seq_along(url_list)){ # 读取当前页面预加载的JSON数据 current_json <- read_html(url_list[n]) |> html_element("script#preload") |> html_text() |> sub("\\s*window\\.__PRELOADED_STATE__ = ", "", x = _, perl = TRUE) |> sub(";\\s*$", "", x = _, perl = TRUE) |> fromJSON(simplifyVector = FALSE) # 提取当前页面对应的路径key current_path <- gsub(".*https://fortune.com","", url_list[n]) current_page_data <- current_json$components$page[[current_path]] # 提取公司信息 current_info <- current_page_data |> find_by_name("body", "children") |> find_by_name("company-about-wrapper", "children") |> find_by_name("company-information", "config") # 存储需要的字段 res_list[[n]] <- data.frame( url = url_list[n], employees = current_info$employees ) # 可选:加2秒延时,避免请求过于频繁被网站封禁 Sys.sleep(2) } # 合并所有结果为数据框 final_result <- do.call(rbind, res_list) # 查看结果 print(final_result)
批量爬取建议
爬取500个页面时建议增加tryCatch容错逻辑,避免单个页面加载失败导致整个循环中断;如果出现部分页面结构不一致的情况,也可以通过容错逻辑跳过异常页面、记录失败url后续补爬。
内容的提问来源于stack exchange,提问作者Xian Zhao
相关产品推荐
相关产品推荐

