You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Rvest抓取MLB阵容时,如何在节点不存在时返回NULL值?

解决MLB阵容抓取时缺失节点导致比赛记录丢失的问题

核心问题原因

你当前使用html_nodes()(复数形式)抓取节点,当阵容未发布时,对应球员节点不存在,html_nodes()会返回空向量。后续用Map()组合数据框时,空向量会导致该场比赛的记录被直接跳过,最终合并数据时无法匹配。

解决方案:用html_node()替代html_nodes()

html_node()是单数形式,它会尝试抓取单个节点:

  • 找到节点时返回对应属性值
  • 找不到节点时返回NA_character_(而非空向量)

这样每个打者位置的变量都会有值(要么是球员ID,要么是NA),确保该场比赛的记录被保留。

修改后的基础代码

# 抓取客队打者ID(找不到则返回NA)
vbatter1 = page %>% html_node(".col--min:nth-child(1) .player:nth-child(1) .player-link") %>% html_attr("data-mlb")
vbatter2 = page %>% html_node(".col--min:nth-child(1) .player:nth-child(2) .player-link") %>% html_attr("data-mlb")
vbatter3 = page %>% html_node(".col--min:nth-child(1) .player:nth-child(3) .player-link") %>% html_attr("data-mlb")
vbatter4 = page %>% html_node(".col--min:nth-child(1) .player:nth-child(4) .player-link") %>% html_attr("data-mlb")
vbatter5 = page %>% html_node(".col--min:nth-child(1) .player:nth-child(5) .player-link") %>% html_attr("data-mlb")
vbatter6 = page %>% html_node(".col--min:nth-child(1) .player:nth-child(6) .player-link") %>% html_attr("data-mlb")
vbatter7 = page %>% html_node(".col--min:nth-child(1) .player:nth-child(7) .player-link") %>% html_attr("data-mlb")
vbatter8 = page %>% html_node(".col--min:nth-child(1) .player:nth-child(8) .player-link") %>% html_attr("data-mlb")
vbatter9 = page %>% html_node(".col--min:nth-child(1) .player:nth-child(9) .player-link") %>% html_attr("data-mlb")

# 抓取主队打者ID(找不到则返回NA)
hbatter1 = page %>% html_node(".col--min+ .col--min .player:nth-child(1) .player-link") %>% html_attr("data-mlb")
hbatter2 = page %>% html_node(".col--min+ .col--min .player:nth-child(2) .player-link") %>% html_attr("data-mlb")
hbatter3 = page %>% html_node(".col--min+ .col--min .player:nth-child(3) .player-link") %>% html_attr("data-mlb")
hbatter4 = page %>% html_node(".col--min+ .col--min .player:nth-child(4) .player-link") %>% html_attr("data-mlb")
hbatter5 = page %>% html_node(".col--min+ .col--min .player:nth-child(5) .player-link") %>% html_attr("data-mlb")
hbatter6 = page %>% html_node(".col--min+ .col--min .player:nth-child(6) .player-link") %>% html_attr("data-mlb")
hbatter7 = page %>% html_node(".col--min+ .col--min .player:nth-child(7) .player-link") %>% html_attr("data-mlb")
hbatter8 = page %>% html_node(".col--min+ .col--min .player:nth-child(8) .player-link") %>% html_attr("data-mlb")
hbatter9 = page %>% html_node(".col--min+ .col--min .player:nth-child(9) .player-link") %>% html_attr("data-mlb")

# 组合数据框,此时缺失阵容的比赛会保留NA值
df <- do.call(rbind, Map(data.frame, GameTime=time, VisTm=VisTm, HmTm=HmTm, VisStPchID=vSP, HmStPchID=hSP, 
                         VisBat1ID=vbatter1, VisBat2ID=vbatter2, VisBat3ID=vbatter3, VisBat4ID=vbatter4, 
                         VisBat5ID=vbatter5, VisBat6ID=vbatter6, VisBat7ID=vbatter7, VisBat8ID=vbatter8, 
                         VisBat9ID=vbatter9, HmBat1ID=hbatter1, HmBat2ID=hbatter2, HmBat3ID=hbatter3, 
                         HmBat4ID=hbatter4, HmBat5ID=hbatter5, HmBat6ID=hbatter6, HmBat7ID=hbatter7, 
                         HmBat8ID=hbatter8, HmBat9ID=hbatter9))

优化:封装函数减少重复代码

如果要抓取多场比赛,重复写9次代码太繁琐,可以封装一个函数批量获取打者ID:

# 加载依赖包
library(rvest)
library(purrr)
library(glue)

# 定义获取单个打者ID的函数
get_batter_id <- function(page, team_selector, batter_position) {
  # 拼接CSS选择器
  selector <- glue("{team_selector} .player:nth-child({batter_position}) .player-link")
  # 抓取节点,找不到则返回NA
  page %>% html_node(selector) %>% html_attr("data-mlb")
}

# 定义客队、主队的基础选择器
vis_team_selector <- ".col--min:nth-child(1)"
hm_team_selector <- ".col--min+ .col--min"

# 批量获取客队9位打者ID
vis_batters <- map_chr(1:9, ~get_batter_id(page, vis_team_selector, .x))
names(vis_batters) <- paste0("VisBat", 1:9, "ID")

# 批量获取主队9位打者ID
hm_batters <- map_chr(1:9, ~get_batter_id(page, hm_team_selector, .x))
names(hm_batters) <- paste0("HmBat", 1:9, "ID")

# 组合单场比赛数据
single_game_data <- c(
  list(GameTime=time, VisTm=VisTm, HmTm=HmTm, VisStPchID=vSP, HmStPchID=hSP),
  as.list(vis_batters),
  as.list(hm_batters)
) %>% data.frame(stringsAsFactors = FALSE)

# 如果是多场比赛,将所有single_game_data存入列表后再合并
df <- do.call(rbind, list_of_single_game_data)

这样处理后,即使阵容未发布,该场比赛的记录也会保留,所有缺失的打者ID字段会填充为NA,完全不影响后续的数据合并匹配。

内容的提问来源于stack exchange,提问作者as21

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 16:32:53