You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R爬取美联储网页遇异常:#article选择器在旧页面无文本返回

问题原因及解决方案

问题原因

美联储网站的页面结构随时间发生了变化:

  • 2023年的会议纪要页面(第一个链接)中,核心文本内容被包裹在id="article"的元素内,所以你的代码能正常提取。
  • 2011年的旧页面(第二个链接)没有id="article"的元素,对应的文本内容在id="content"下的.container容器中,因此用#article选择器无法匹配到任何内容,返回空白。

解决方案

你可以编写一个兼容不同页面结构的爬虫逻辑,先尝试匹配新页面的选择器,失败后再切换到旧页面的选择器。示例代码如下:

library(rvest)
library(stringr)

# 定义通用爬取函数
scrape_fomc_minutes <- function(link) {
  page <- read_html(link)
  
  # 优先尝试新页面选择器
  text_content <- page %>% html_nodes("#article") %>% html_text()
  
  # 如果未获取到内容,切换到旧页面选择器
  if (length(text_content) == 0) {
    text_content <- page %>% html_nodes("#content .container") %>% html_text()
  }
  
  # 合并文本内容
  text_content <- str_c(text_content, collapse = " ")
  return(text_content)
}

# 测试两个链接
link_2023 <- "https://www.federalreserve.gov/monetarypolicy/fomcminutes20230201.htm"
link_2011 <- "https://www.federalreserve.gov/monetarypolicy/fomcminutes20111102.htm"

# 调用函数获取文本
scrape_fomc_minutes(link_2023)
scrape_fomc_minutes(link_2011)

验证方法

你可以通过浏览器的开发者工具(按F12)查看页面结构:

  • 打开目标网页后,使用"选择元素"工具定位到纪要文本,查看其所在的父容器的id或class属性,以此确定正确的CSS选择器。

内容的提问来源于stack exchange,提问作者James Rider

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 13:54:56