You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用rvest爬取Sports-Reference球队数据失败求助

解决rvest爬取Sports Reference球队数据的问题

Sports Reference页面的team_stats表格数据被包裹在HTML注释中,直接使用rvest::html_table()无法抓取到——因为默认解析逻辑会忽略注释内的内容。以下是具体解决步骤:

核心代码实现

library(rvest)
library(xml2)

# 目标页面URL
url <- "https://www.sports-reference.com/cfb/boxscores/2022-09-02-charlotte.html"

# 读取页面内容
page <- read_html(url)

# 定位包含team_stats的HTML注释块并转换为可解析的HTML
team_stats_comment <- page %>%
  html_nodes(xpath = "//comment()[contains(., 'id=\"team_stats\"')]") %>%
  html_text() %>%
  read_html()

# 提取并整理球队数据表格
team_stats <- team_stats_comment %>%
  html_node("#team_stats") %>%
  html_table(header = TRUE, fill = TRUE)

# 清理表格内重复的表头行(可选,根据实际返回结果调整)
team_stats <- team_stats %>%
  dplyr::filter(!dplyr::row_number() %in% which(.[[1]] == "Team"))

关键说明

  • 页面注释块通过XPath定位://comment()[contains(., 'id=\"team_stats\"')] 精准匹配包含目标表格ID的注释内容
  • 将注释文本转换为HTML节点后,即可用常规rvest方法提取表格
  • 部分返回结果会包含重复表头行,可通过过滤逻辑清除

内容的提问来源于stack exchange,提问作者Nathan Arizona

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 00:50:35