You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用rvest多页爬取时提取href属性返回NA生成错误链接问题求助

问题原因
  • 直接原因:你选择的节点td.views-field-title是表格单元格节点,href属性属于单元格内部的<a>标签,直接从单元格提取href会返回空值(NA),拼接域名后就生成了无效链接。
  • 附带问题1:当前使用的数据源为劳拉·布什的演讲归档页,若要提取梅拉尼娅·特朗普的演讲,需要替换为对应的归档页面。
  • 附带问题2:原有代码中的get_text函数存在两个错误:一是函数内误用了全局变量title_links而非传入的参数title_link,二是返回值错误设置为页面对象而非爬取到的演讲文本,后续运行也会报错。
修正后的可用代码
library(rvest)
library(dplyr)

# 注意:当前链接为劳拉·布什的演讲页,需替换为梅拉尼娅·特朗普的对应归档页后再运行爬取目标内容
link <- "https://www.presidency.ucsb.edu/documents/presidential-documents-archive-guidebook/remarks-and-statements-the-first-lady-laura-bush"
page <- read_html(link)

# 提取演讲标题、链接
title <- page %>% html_nodes("td.views-field-title a") %>% html_text()
title_links <- page %>% html_nodes("td.views-field-title a") %>%
  html_attr("href") %>% paste0("https://www.presidency.ucsb.edu", .)

# 提取演讲年份
year <- page %>% html_nodes(".date-display-single") %>% html_text()

# 提取演讲人姓名
flotus <- page %>% html_nodes(".views-field-title-1.nowrap") %>% html_text()

# 修正后的单页内容爬取函数
get_text <- function(title_link){
  speech_page <- read_html(title_link)
  speech_text <- speech_page %>% 
    html_nodes(".field-docs-content p") %>%
    html_text() %>% 
    paste(collapse = "\n")
  return(speech_text)
}

# 批量爬取内容,加间隔避免触发反爬机制
text <- sapply(title_links, function(x) {
  Sys.sleep(1)
  get_text(x)
})

# 合并为结构化数据框可直接导出
result <- data.frame(
  title = title,
  year = year,
  speaker = flotus,
  link = title_links,
  content = text,
  stringsAsFactors = FALSE
)

内容的提问来源于stack exchange,提问作者w5698

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 11:06:03