You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用rvest与purrr爬取多页嵌套链接时如何提取对应演讲正文

实现方案

你需要先从列表页提取每篇演讲的详情页链接,再逐页抓取正文内容即可,完整修改后的可运行代码和修改说明如下:


完整修改后代码
library(rvest)
library(purrr)
library(dplyr)

# 列表页基础URL
url_base <- "https://www.presidency.ucsb.edu/documents/presidential-documents-archive-guidebook/remarks-and-statements-the-first-lady-laura-bush?page=%d"
# 站点域名,用于补全详情页相对路径
domain <- "https://www.presidency.ucsb.edu"

# 第一步:抓取所有列表页的基础信息+详情页链接
map_df(0:16, function(i) { 
  # 分页从0开始,修正原代码1:17漏掉第一页的问题
  cat(".")
  pg <- read_html(sprintf(url_base, i))
  
  # 提取标题节点,同时拿文本和链接
  title_nodes <- html_nodes(pg, "td.views-field-title a")
  
  data.frame(
    name = html_text(html_nodes(pg, ".views-field-title-1.nowrap")),
    title = html_text(title_nodes),
    page_url = paste0(domain, html_attr(title_nodes, "href")),
    year = html_text(html_nodes(pg, ".date-display-single")),
    stringsAsFactors = FALSE
  )
}) -> flotus_base

# 第二步:定义单篇演讲正文爬取函数
get_speech_content <- function(url) {
  cat("+")
  # 每次请求暂停1秒,避免触发反爬规则
  Sys.sleep(1)
  # 错误处理:单页请求失败返回NA,不中断整体爬取
  tryCatch({
    content_pg <- read_html(url)
    # 提取正文并清理多余空白
    content <- html_text(html_nodes(content_pg, ".field-docs-content")) %>% trimws()
    return(content)
  }, error = function(e) {
    return(NA_character_)
  })
}

# 第三步:批量抓取所有演讲正文,合并到原数据框
flotus <- flotus_base %>%
  mutate(
    speech_content = map_chr(page_url, get_speech_content)
  )

关键修改说明
  • 修正了分页参数:目标站点分页从page=0对应第一页,原代码的1:17会漏掉第一页数据,改为0:16覆盖全部17页内容
  • 新增详情页链接抓取:从列表页的标题锚点中提取相对路径,拼接域名得到完整的演讲详情页访问地址
  • 新增正文爬取逻辑:单独封装了正文爬取函数,加入请求延迟和错误处理,保证爬取稳定性
  • 正文选择器用站点统一的正文容器类.field-docs-content,提取后做了空白字符清理

内容的提问来源于stack exchange,提问作者w5698

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 09:45:04