如何通过循环批量处理URL实现多页面网页爬取(R语言)
批量处理R爬虫URL的循环实现方法
你的核心需求是通过批量生成URL并自动调用爬虫函数,避免重复编写单个链接和调用代码。以下是几种高效的实现方式:
1. 批量生成URL
观察你的URL可知,只有Pagina参数从1到10变化,因此可以用paste0或glue快速生成所有页面的链接:
用paste0生成(无需额外包)
# 生成1到10页的URL向量 urls <- paste0( "https://santabarbara.siscam.com.br/Documentos/Pesquisa/74?Pesquisa=Simples&Pagina=", 1:10, "&Documento=117&Modulo=8&AnoInicial=2022" )
用glue生成(更易读)
如果安装了glue包,写法更直观:
library(glue) urls <- glue( "https://santabarbara.siscam.com.br/Documentos/Pesquisa/74?Pesquisa=Simples&Pagina={1:10}&Documento=117&Modulo=8&AnoInicial=2022" )
2. 批量调用爬虫函数
方法一:用lapply实现(简洁高效)
lapply会遍历URL向量,逐个调用scraper函数,返回结果列表,最后用do.call(rbind, ...)合并为单个数据框:
library(rvest) library(dplyr) # 保留你已写好的scraper函数 scraper <- function(link){ page = read_html(link) titulo = page %>% html_nodes("h4 a") %>% html_text() tipo = page %>% html_nodes("h4+ .row .col-md-4") %>% html_text() data = page %>% html_nodes("p.col-md-6") %>% html_text() protocolo = page %>% html_nodes(".row:nth-child(3) .col-md-4") %>% html_text() situacao = page %>% html_nodes(".row~ .row+ .row p.col-md-4:nth-child(1)") %>% html_text() regime = page %>% html_nodes("p.col-md-4:nth-child(2)") %>% html_text() quorum = page %>% html_nodes(".col-md-4~ .col-md-4+ .col-md-4") %>% html_text() autoria = page %>% html_nodes(".row:nth-child(5) .col-md-12") %>% html_text() assunto = page %>% html_nodes(".row:nth-child(6) .col-md-12") %>% html_text() result <- data.frame(titulo, tipo, data, protocolo, situacao, regime, quorum, autoria, assunto) return(result) } # 批量爬取并合并结果 results_list <- lapply(urls, scraper) final_result <- do.call(rbind, results_list)
方法二:用for循环实现(逻辑清晰)
如果你更习惯循环逻辑,也可以用for循环逐个处理:
# 初始化空列表存储每页结果 results_list <- list() # 循环处理1到10页 for(i in 1:10){ # 生成当前页URL current_url <- paste0( "https://santabarbara.siscam.com.br/Documentos/Pesquisa/74?Pesquisa=Simples&Pagina=", i, "&Documento=117&Modulo=8&AnoInicial=2022" ) # 调用爬虫并存储结果 results_list[[i]] <- scraper(current_url) } # 合并所有结果为单个数据框 final_result <- do.call(rbind, results_list)
3. 可选优化:添加错误处理
为避免单个页面爬取失败导致整个流程中断,建议给scraper函数加上tryCatch错误捕获:
scraper <- function(link){ tryCatch({ page = read_html(link) titulo = page %>% html_nodes("h4 a") %>% html_text() tipo = page %>% html_nodes("h4+ .row .col-md-4") %>% html_text() data = page %>% html_nodes("p.col-md-6") %>% html_text() protocolo = page %>% html_nodes(".row:nth-child(3) .col-md-4") %>% html_text() situacao = page %>% html_nodes(".row~ .row+ .row p.col-md-4:nth-child(1)") %>% html_text() regime = page %>% html_nodes("p.col-md-4:nth-child(2)") %>% html_text() quorum = page %>% html_nodes(".col-md-4~ .col-md-4+ .col-md-4") %>% html_text() autoria = page %>% html_nodes(".row:nth-child(5) .col-md-12") %>% html_text() assunto = page %>% html_nodes(".row:nth-child(6) .col-md-12") %>% html_text() result <- data.frame(titulo, tipo, data, protocolo, situacao, regime, quorum, autoria, assunto) return(result) }, error = function(e){ # 打印错误信息,方便排查 message(paste("爬取链接", link, "失败:", e$message)) # 返回空数据框,不影响后续合并 return(data.frame()) }) }
内容的提问来源于stack exchange,提问作者Lucas Esteves
相关产品推荐
相关产品推荐

