You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用rvest抓取baseball-reference.com2018球员标准击球表遇阻:仅获首个表

解决rvest抓取baseball-reference目标表格的问题

我懂你的困扰——你想抓取2018年MLB「Player Standard Batting」表格,但当前代码只拿到了页面的第一个表,对吧?这是因为baseball-reference为了限制批量抓取,把这类核心数据表格藏在了HTML注释(<!-- ... -->)里,常规的html_table()不会自动解析注释内的内容。

下面是修正后的代码,专门定位到目标表格所在的注释节点,帮你拿到想要的数据:

library(rvest)
url <- "https://www.baseball-reference.com/leagues/MLB/2018-standard-batting.shtml"
pg <- read_html(url)

# 定位包含目标表格的注释节点
target_comment <- pg %>% 
  html_nodes(xpath = '//comment()') %>% 
  # 筛选出带有目标表格标识的注释
  keep(~grepl("Player Standard Batting", .x %>% html_text())) %>% 
  first()

# 将注释内容转换为可解析的HTML对象
comment_html <- target_comment %>% 
  html_text() %>% 
  read_html()

# 提取你需要的目标表格
df_batting <- comment_html %>% 
  html_table(fill = TRUE) %>% 
  first()

代码说明:

  • 先用html_nodes(xpath = '//comment()')抓取页面所有注释节点,再筛选出包含「Player Standard Batting」的目标注释。
  • 把注释里的文本内容转为HTML对象,这样就能用html_table()正常解析表格了。
  • 最后提取的表格就是你要找的第二个目标表。

运行这段代码应该就能得到你需要的击球数据了。

内容的提问来源于stack exchange,提问作者JAGIMAC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:45:37