You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用rvest爬取网页返回空xml_nodeset问题求助

解决rvest爬取头盔评级网站返回{xml_nodeset (0)}的问题

你遇到的问题是因为span.helmet-name是浏览器通过JavaScript动态生成的元素,而read_html()只能获取网页的静态原始HTML,静态源码里不存在这个类的元素,所以无法匹配到。

解决方案1:直接抓取静态HTML中的头盔名称

查看目标网页的静态源码,头盔名称实际位于表格第一列的<b>标签内,用以下代码即可提取:

library(rvest)
library(dplyr)

url <- "https://www.helmet.beam.vt.edu/bicycle-helmet-ratings.html"
webpage <- read_html(url)

# 提取所有头盔名称
helmet_names <- webpage %>% 
  html_nodes("table#ratings-table tr td:nth-child(1) b") %>% 
  html_text()

# 查看前5个结果
head(helmet_names, 5)

解决方案2:处理动态渲染内容(若后续需要)

如果后续遇到完全依赖JS加载的内容,可以用RSelenium模拟浏览器渲染页面,再提取元素:

library(RSelenium)
library(rvest)

# 启动Chrome驱动(需提前安装ChromeDriver并配置环境变量)
driver <- rsDriver(browser = "chrome", port = 4567L)
remDr <- driver[["client"]]

# 访问目标页面
remDr$navigate("https://www.helmet.beam.vt.edu/bicycle-helmet-ratings.html")

# 获取渲染后的页面源码
page_source <- remDr$getPageSource()[[1]]
webpage <- read_html(page_source)

# 用你指定的选择器提取头盔名称
helmet_names <- webpage %>% 
  html_nodes("span.helmet-name") %>% 
  html_text()

head(helmet_names, 5)

# 关闭浏览器和驱动
remDr$close()
driver$server$stop()

优先推荐方案1,因为它不需要额外依赖,执行效率更高,且完全能满足当前需求。

内容的提问来源于stack exchange,提问作者Andrew Jackson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 15:22:09