使用Rvest抓取MLB阵容时,如何在节点不存在时返回NULL值?
解决MLB阵容抓取时缺失节点导致比赛记录丢失的问题
核心问题原因
你当前使用html_nodes()(复数形式)抓取节点,当阵容未发布时,对应球员节点不存在,html_nodes()会返回空向量。后续用Map()组合数据框时,空向量会导致该场比赛的记录被直接跳过,最终合并数据时无法匹配。
解决方案:用html_node()替代html_nodes()
html_node()是单数形式,它会尝试抓取单个节点:
- 找到节点时返回对应属性值
- 找不到节点时返回
NA_character_(而非空向量)
这样每个打者位置的变量都会有值(要么是球员ID,要么是NA),确保该场比赛的记录被保留。
修改后的基础代码
# 抓取客队打者ID(找不到则返回NA) vbatter1 = page %>% html_node(".col--min:nth-child(1) .player:nth-child(1) .player-link") %>% html_attr("data-mlb") vbatter2 = page %>% html_node(".col--min:nth-child(1) .player:nth-child(2) .player-link") %>% html_attr("data-mlb") vbatter3 = page %>% html_node(".col--min:nth-child(1) .player:nth-child(3) .player-link") %>% html_attr("data-mlb") vbatter4 = page %>% html_node(".col--min:nth-child(1) .player:nth-child(4) .player-link") %>% html_attr("data-mlb") vbatter5 = page %>% html_node(".col--min:nth-child(1) .player:nth-child(5) .player-link") %>% html_attr("data-mlb") vbatter6 = page %>% html_node(".col--min:nth-child(1) .player:nth-child(6) .player-link") %>% html_attr("data-mlb") vbatter7 = page %>% html_node(".col--min:nth-child(1) .player:nth-child(7) .player-link") %>% html_attr("data-mlb") vbatter8 = page %>% html_node(".col--min:nth-child(1) .player:nth-child(8) .player-link") %>% html_attr("data-mlb") vbatter9 = page %>% html_node(".col--min:nth-child(1) .player:nth-child(9) .player-link") %>% html_attr("data-mlb") # 抓取主队打者ID(找不到则返回NA) hbatter1 = page %>% html_node(".col--min+ .col--min .player:nth-child(1) .player-link") %>% html_attr("data-mlb") hbatter2 = page %>% html_node(".col--min+ .col--min .player:nth-child(2) .player-link") %>% html_attr("data-mlb") hbatter3 = page %>% html_node(".col--min+ .col--min .player:nth-child(3) .player-link") %>% html_attr("data-mlb") hbatter4 = page %>% html_node(".col--min+ .col--min .player:nth-child(4) .player-link") %>% html_attr("data-mlb") hbatter5 = page %>% html_node(".col--min+ .col--min .player:nth-child(5) .player-link") %>% html_attr("data-mlb") hbatter6 = page %>% html_node(".col--min+ .col--min .player:nth-child(6) .player-link") %>% html_attr("data-mlb") hbatter7 = page %>% html_node(".col--min+ .col--min .player:nth-child(7) .player-link") %>% html_attr("data-mlb") hbatter8 = page %>% html_node(".col--min+ .col--min .player:nth-child(8) .player-link") %>% html_attr("data-mlb") hbatter9 = page %>% html_node(".col--min+ .col--min .player:nth-child(9) .player-link") %>% html_attr("data-mlb") # 组合数据框,此时缺失阵容的比赛会保留NA值 df <- do.call(rbind, Map(data.frame, GameTime=time, VisTm=VisTm, HmTm=HmTm, VisStPchID=vSP, HmStPchID=hSP, VisBat1ID=vbatter1, VisBat2ID=vbatter2, VisBat3ID=vbatter3, VisBat4ID=vbatter4, VisBat5ID=vbatter5, VisBat6ID=vbatter6, VisBat7ID=vbatter7, VisBat8ID=vbatter8, VisBat9ID=vbatter9, HmBat1ID=hbatter1, HmBat2ID=hbatter2, HmBat3ID=hbatter3, HmBat4ID=hbatter4, HmBat5ID=hbatter5, HmBat6ID=hbatter6, HmBat7ID=hbatter7, HmBat8ID=hbatter8, HmBat9ID=hbatter9))
优化:封装函数减少重复代码
如果要抓取多场比赛,重复写9次代码太繁琐,可以封装一个函数批量获取打者ID:
# 加载依赖包 library(rvest) library(purrr) library(glue) # 定义获取单个打者ID的函数 get_batter_id <- function(page, team_selector, batter_position) { # 拼接CSS选择器 selector <- glue("{team_selector} .player:nth-child({batter_position}) .player-link") # 抓取节点,找不到则返回NA page %>% html_node(selector) %>% html_attr("data-mlb") } # 定义客队、主队的基础选择器 vis_team_selector <- ".col--min:nth-child(1)" hm_team_selector <- ".col--min+ .col--min" # 批量获取客队9位打者ID vis_batters <- map_chr(1:9, ~get_batter_id(page, vis_team_selector, .x)) names(vis_batters) <- paste0("VisBat", 1:9, "ID") # 批量获取主队9位打者ID hm_batters <- map_chr(1:9, ~get_batter_id(page, hm_team_selector, .x)) names(hm_batters) <- paste0("HmBat", 1:9, "ID") # 组合单场比赛数据 single_game_data <- c( list(GameTime=time, VisTm=VisTm, HmTm=HmTm, VisStPchID=vSP, HmStPchID=hSP), as.list(vis_batters), as.list(hm_batters) ) %>% data.frame(stringsAsFactors = FALSE) # 如果是多场比赛,将所有single_game_data存入列表后再合并 df <- do.call(rbind, list_of_single_game_data)
这样处理后,即使阵容未发布,该场比赛的记录也会保留,所有缺失的打者ID字段会填充为NA,完全不影响后续的数据合并匹配。
内容的提问来源于stack exchange,提问作者as21
相关产品推荐
相关产品推荐

