You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过循环批量处理URL实现多页面网页爬取(R语言)

批量处理R爬虫URL的循环实现方法

你的核心需求是通过批量生成URL并自动调用爬虫函数,避免重复编写单个链接和调用代码。以下是几种高效的实现方式:

1. 批量生成URL

观察你的URL可知,只有Pagina参数从1到10变化,因此可以用paste0或glue快速生成所有页面的链接:

用paste0生成(无需额外包)

# 生成1到10页的URL向量
urls <- paste0(
  "https://santabarbara.siscam.com.br/Documentos/Pesquisa/74?Pesquisa=Simples&Pagina=",
  1:10,
  "&Documento=117&Modulo=8&AnoInicial=2022"
)

用glue生成(更易读)

如果安装了glue包,写法更直观:

library(glue)
urls <- glue(
  "https://santabarbara.siscam.com.br/Documentos/Pesquisa/74?Pesquisa=Simples&Pagina={1:10}&Documento=117&Modulo=8&AnoInicial=2022"
)

2. 批量调用爬虫函数

方法一:用lapply实现(简洁高效)

lapply会遍历URL向量,逐个调用scraper函数,返回结果列表,最后用do.call(rbind, ...)合并为单个数据框:

library(rvest)
library(dplyr)

# 保留你已写好的scraper函数
scraper <- function(link){
  page = read_html(link)
  titulo = page %>% html_nodes("h4 a") %>% html_text()
  tipo = page %>% html_nodes("h4+ .row .col-md-4") %>% html_text()
  data = page %>% html_nodes("p.col-md-6") %>% html_text()
  protocolo = page %>% html_nodes(".row:nth-child(3) .col-md-4") %>% html_text()
  situacao = page %>% html_nodes(".row~ .row+ .row p.col-md-4:nth-child(1)") %>% html_text()
  regime = page %>% html_nodes("p.col-md-4:nth-child(2)") %>% html_text()
  quorum = page %>% html_nodes(".col-md-4~ .col-md-4+ .col-md-4") %>% html_text()
  autoria = page %>% html_nodes(".row:nth-child(5) .col-md-12") %>% html_text()
  assunto = page %>% html_nodes(".row:nth-child(6) .col-md-12") %>% html_text()
  
  result <- data.frame(titulo, tipo, data, protocolo, situacao, regime, quorum, autoria, assunto)
  
  return(result)
}

# 批量爬取并合并结果
results_list <- lapply(urls, scraper)
final_result <- do.call(rbind, results_list)

方法二:用for循环实现(逻辑清晰)

如果你更习惯循环逻辑,也可以用for循环逐个处理:

# 初始化空列表存储每页结果
results_list <- list()

# 循环处理1到10页
for(i in 1:10){
  # 生成当前页URL
  current_url <- paste0(
    "https://santabarbara.siscam.com.br/Documentos/Pesquisa/74?Pesquisa=Simples&Pagina=",
    i,
    "&Documento=117&Modulo=8&AnoInicial=2022"
  )
  # 调用爬虫并存储结果
  results_list[[i]] <- scraper(current_url)
}

# 合并所有结果为单个数据框
final_result <- do.call(rbind, results_list)

3. 可选优化:添加错误处理

为避免单个页面爬取失败导致整个流程中断,建议给scraper函数加上tryCatch错误捕获:

scraper <- function(link){
  tryCatch({
    page = read_html(link)
    titulo = page %>% html_nodes("h4 a") %>% html_text()
    tipo = page %>% html_nodes("h4+ .row .col-md-4") %>% html_text()
    data = page %>% html_nodes("p.col-md-6") %>% html_text()
    protocolo = page %>% html_nodes(".row:nth-child(3) .col-md-4") %>% html_text()
    situacao = page %>% html_nodes(".row~ .row+ .row p.col-md-4:nth-child(1)") %>% html_text()
    regime = page %>% html_nodes("p.col-md-4:nth-child(2)") %>% html_text()
    quorum = page %>% html_nodes(".col-md-4~ .col-md-4+ .col-md-4") %>% html_text()
    autoria = page %>% html_nodes(".row:nth-child(5) .col-md-12") %>% html_text()
    assunto = page %>% html_nodes(".row:nth-child(6) .col-md-12") %>% html_text()
    
    result <- data.frame(titulo, tipo, data, protocolo, situacao, regime, quorum, autoria, assunto)
    
    return(result)
  }, error = function(e){
    # 打印错误信息,方便排查
    message(paste("爬取链接", link, "失败:", e$message))
    # 返回空数据框,不影响后续合并
    return(data.frame())
  })
}

内容的提问来源于stack exchange,提问作者Lucas Esteves

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 15:10:41