You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Rvest爬取Europol新闻页遇read_html空对象问题求助

解决Europol新闻室页面爬取无内容问题

问题概况

  • 爬取目标:https://www.europol.europa.eu/media-press/newsroom
  • 遇到的问题:
    • 直接用read_html()读取页面,返回近乎空的嵌套列表,无有效内容
    • 尝试RSelenium模拟浏览器后,使用div.content-wrapper、h3 a等CSS选择器/Xpath配合html_nodes()/html_elements(),始终返回character(0)
  • 用户测试代码:
read_html("https://www.europol.europa.eu/media-press/newsroom") %>% 
  html_nodes("div.content-wrapper") %>%
  html_attr("href")

核心原因

该页面的新闻内容依赖JavaScript动态渲染,且可能存在基础反爬检测(比如验证请求是否来自真实浏览器)。read_html()仅能抓取静态HTML,无法触发JS加载;RSelenium若未配置等待逻辑,也会在内容未加载完成时就获取源码,导致选择器匹配失败。

具体解决方法

1. 优化RSelenium配置,确保页面完全加载

通过等待JS执行、模拟滚动触发内容加载,获取完整渲染后的页面:

library(RSelenium)
library(rvest)

# 启动Chrome驱动(需提前安装对应版本的ChromeDriver)
driver <- rsDriver(browser = "chrome", port = 4567L)
remDr <- driver[["client"]]

# 访问目标页面
remDr$navigate("https://www.europol.europa.eu/media-press/newsroom")

# 等待页面初始加载(时间可根据网络情况调整)
Sys.sleep(5)

# 滚动页面触发更多内容加载(若新闻为滚动加载模式)
remDr$executeScript("window.scrollTo(0, document.body.scrollHeight);")
Sys.sleep(3)

# 获取渲染完成的页面源码
page_source <- remDr$getPageSource()[[1]]
html <- read_html(page_source)

# 提取新闻链接(以实际页面的h3下a标签为例)
news_links <- html %>% 
  html_nodes("h3 a") %>% 
  html_attr("href")

# 关闭驱动
remDr$close()
driver$server$stop()

# 查看结果
print(news_links)

2. 模拟浏览器请求头(备选方案)

若不想使用RSelenium,可通过httr添加完整浏览器请求头,尝试绕过基础反爬:

library(httr)
library(rvest)

# 模拟Chrome浏览器请求头
headers <- c(
  "User-Agent" = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
  "Accept" = "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8",
  "Accept-Language" = "en-US,en;q=0.5",
  "Connection" = "keep-alive"
)

# 发送请求并解析
response <- GET("https://www.europol.europa.eu/media-press/newsroom", add_headers(.headers = headers))
html <- read_html(content(response, "text"))

# 提取内容
news_links <- html %>% 
  html_nodes("h3 a") %>% 
  html_attr("href")

print(news_links)

3. 验证选择器正确性

打开页面开发者工具(F12),在Elements面板中搜索目标选择器:

  • 右键目标元素→复制→复制CSS选择器/Xpath,确保使用的选择器与页面实际元素匹配,避免因类名、结构变化导致匹配失败。

内容的提问来源于stack exchange,提问作者Matis Poussardin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 10:00:28