You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用rvest的XPath无法抓取非活跃球队表数据求助

解决非活跃球队数据抓取问题

问题原因

非活跃球队的表格内容被页面以HTML注释的形式包裹,直接使用常规XPath选择器无法定位到目标元素,这是你之前尝试失败的核心原因。

完整解决方案代码

library(tidyverse)
library(rvest)

teams_page <- 'https://www.pro-football-reference.com/teams/'

# 抓取活跃球队三字母缩写
active_teams <- read_html(teams_page) %>% 
  html_elements(xpath = '//*[@id="teams_active"]/tbody//a[@href]') %>% 
  html_attr("href") %>% 
  str_extract("(?<=teams/)[A-Z]{3}(?=/)") %>% 
  na.omit()

# 抓取非活跃球队三字母缩写
inactive_teams <- read_html(teams_page) %>% 
  # 定位包含非活跃表格的div
  html_element(xpath = '//*[@id="all_teams_inactive"]') %>% 
  # 提取其中的HTML注释内容
  html_text() %>% 
  # 将注释内容转为可解析的HTML
  read_html() %>% 
  # 定位非活跃表格的tbody下的链接
  html_elements(xpath = '//*[@id="teams_inactive"]/tbody//a[@href]') %>% 
  html_attr("href") %>% 
  str_extract("(?<=teams/)[A-Z]{3}(?=/)") %>% 
  na.omit()

# 合并结果
all_teams_abbr <- c(active_teams, inactive_teams)
print(all_teams_abbr)

代码说明

  • 活跃球队的抓取逻辑优化了XPath写法,直接用str_extract提取三字母缩写,更简洁高效。
  • 非活跃球队部分:
    1. 先定位到包裹注释的div#all_teams_inactive容器
    2. 提取容器内的文本(即被注释的HTML内容)
    3. 将注释内容重新解析为可操作的HTML对象
    4. 后续逻辑与活跃球队一致,精准提取链接中的三字母缩写

内容的提问来源于stack exchange,提问作者Jeff Henderson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 03:46:13