You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R编写网页爬取函数?多页爬取及函数无输出问题求助

问题解决与多页爬取指导

一、修复函数无输出的问题

你的爬取函数核心逻辑没问题,但没有主动返回结果——R函数默认不会输出赋值语句的结果,需要明确返回数据框。

修改后的函数:

library(rvest)
library(dplyr)

scraper <- function(link){
  page = read_html(link)
  titulo = page %>% html_nodes("h4 a") %>% html_text()
  tipo = page %>% html_nodes("h4+ .row .col-md-4") %>% html_text()
  data = page %>% html_nodes("p.col-md-6") %>% html_text()
  protocolo = page %>% html_nodes(".row:nth-child(3) .col-md-4") %>% html_text()
  situacao = page %>% html_nodes(".row~ .row+ .row p.col-md-4:nth-child(1)") %>% html_text()
  regime = page %>% html_nodes("p.col-md-4:nth-child(2)") %>% html_text()
  quorum = page %>% html_nodes(".col-md-4~ .col-md-4+ .col-md-4") %>% html_text()
  autoria = page %>% html_nodes(".row:nth-child(5) .col-md-12") %>% html_text()
  assunto = page %>% html_nodes(".row:nth-child(6) .col-md-12") %>% html_text()
  
  # 直接返回数据框,无需额外赋值后再return
  data.frame(titulo, tipo, data, protocolo, situacao, regime, quorum, autoria, assunto)
}

调用scraper(link)即可得到当前页的爬取结果。

二、实现10页批量爬取

观察你的链接,Pagina=1是页码参数,只需循环替换该参数值,即可批量爬取所有页面。

方法1:for循环实现(直观易理解)

# 定义基础链接,用%d作为页码占位符
base_link <- "https://santabarbara.siscam.com.br/Documentos/Pesquisa/74?Pesquisa=Simples&Pagina=%d&Documento=117&Modulo=8&AnoInicial=2022"

# 初始化空数据框存储所有结果
all_data <- data.frame()

# 循环爬取1-10页
for (page_num in 1:10) {
  # 生成当前页的完整链接
  current_link <- sprintf(base_link, page_num)
  # 爬取当前页数据
  page_result <- scraper(current_link)
  # 合并到总数据框
  all_data <- rbind(all_data, page_result)
  # 延迟1秒,避免请求过于频繁被服务器封禁
  Sys.sleep(1)
}

# 查看最终结果
print(all_data)

方法2:purrr包批量处理(代码更简洁)

library(purrr)

base_link <- "https://santabarbara.siscam.com.br/Documentos/Pesquisa/74?Pesquisa=Simples&Pagina=%d&Documento=117&Modulo=8&AnoInicial=2022"

# 生成10页的所有链接
all_links <- sprintf(base_link, 1:10)

# 批量爬取并自动合并结果
all_data <- map_dfr(all_links, function(link) {
  Sys.sleep(1)
  scraper(link)
})

print(all_data)

注意事项

  • 必须提前加载依赖包:rvest(网页解析)、dplyr(管道操作),使用purrr时需额外加载。
  • Sys.sleep(1)是强制延迟,可根据服务器响应情况调整时长,避免IP被封禁。
  • 如果某一页没有数据,爬取函数会返回空数据框,合并时不会报错,可根据需求添加空页判断逻辑。

内容的提问来源于stack exchange,提问作者Lucas Esteves

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 14:05:20