如何使用R的rvest库抓取Basketball Reference的2021年NBA季后赛扩展排名表
可以实现,你之前抓取失败的核心原因是该站点的Expanded Standings表格默认被包裹在HTML注释标签中,rvest直接读取节点时会跳过注释内容,所以找不到对应表格。
以下是可用的抓取代码:
# 加载依赖包 library(rvest) library(dplyr) url <- "https://www.basketball-reference.com/playoffs/NBA_2021_standings.html" # 读取页面原始内容 page <- read_html(url) # 提取页面所有注释节点,筛选出包含目标表格的注释段重新解析 expanded_table <- page %>% html_nodes(xpath = "//comment()") %>% html_text() %>% .[grepl("id=\"expanded_standings\"", .)] %>% read_html() %>% html_node("table#expanded_standings") %>% html_table(header = TRUE, fill = TRUE) # 清理冗余表头 expanded_table <- expanded_table[-1,] colnames(expanded_table) <- gsub("\\s+", "_", colnames(expanded_table))
执行完成后expanded_table就是完整的Expanded Standings表数据,可直接在RStudio中做后续分析。
内容的提问来源于stack exchange,提问作者melish
相关产品推荐
相关产品推荐

