如何在R中遍历比赛ID列表并将结果合并至DataFrame?
解决NHL比赛数据爬取合并问题
问题原因
你当前的代码每次循环都会用新比赛的结果覆盖output变量,循环结束后仅保留最后一场的数据,没有累积所有比赛的结果。
解决方案一:循环累积 + 最终合并
先初始化空列表存储每场比赛的数据,循环时将单场结果追加到列表,最后统一合并。同时建议添加game_id字段,方便区分不同比赛的数据:
library(rvest) library(dplyr) library(purrr) library(tidyr) game_ids <- list(401559239,401559240) url_ = "https://www.espn.com/nhl/boxscore/_/gameId" # 初始化空列表存储每场比赛结果 all_results <- list() for (game_id in game_ids) { url2 = paste(url_, game_id, sep = '/') boxscore <- read_html(url2) %>% html_elements("div.Boxscore div.Wrapper") %>% set_names(html_elements(., ".BoxscoreItem__TeamName") %>% html_text()) %>% map(\(team_section) html_elements(team_section, "table")) %>% map(\(team_tables) list( tbl_1 = html_table(team_tables[1:2]) %>% bind_cols(.name_repair = "minimal") %>% set_names(.[1,]) %>% rename(player = Skaters) %>% mutate(position = if_else(G == "G", player, NA), .before = 1) %>% fill(position, .direction = "down") %>% filter(G != "G"), tbl_2 = html_table(team_tables[3:4]) %>% bind_cols(.name_repair = "minimal") %>% set_names(.[1,]) %>% filter(SA != "SA") )) # 提取当前比赛的skater数据,添加game_id字段 current_output <- boxscore %>% map("tbl_1") %>% list_rbind(names_to = "team") %>% mutate(game_id = game_id, .before = 1) # 添加比赛ID # 将当前结果追加到列表 all_results[[length(all_results) + 1]] <- current_output } # 合并所有比赛数据为一个DataFrame final_df <- list_rbind(all_results)
解决方案二:用purrr::map_dfr简化代码
利用tidyverse的map_dfr函数,直接对game_ids列表映射处理逻辑,自动将结果合并为DataFrame,代码更简洁:
library(rvest) library(dplyr) library(purrr) library(tidyr) game_ids <- list(401559239,401559240) url_ = "https://www.espn.com/nhl/boxscore/_/gameId" # 定义处理单场比赛的函数 process_game <- function(game_id) { url2 = paste(url_, game_id, sep = '/') boxscore <- read_html(url2) %>% html_elements("div.Boxscore div.Wrapper") %>% set_names(html_elements(., ".BoxscoreItem__TeamName") %>% html_text()) %>% map(\(team_section) html_elements(team_section, "table")) %>% map(\(team_tables) list( tbl_1 = html_table(team_tables[1:2]) %>% bind_cols(.name_repair = "minimal") %>% set_names(.[1,]) %>% rename(player = Skaters) %>% mutate(position = if_else(G == "G", player, NA), .before = 1) %>% fill(position, .direction = "down") %>% filter(G != "G"), tbl_2 = html_table(team_tables[3:4]) %>% bind_cols(.name_repair = "minimal") %>% set_names(.[1,]) %>% filter(SA != "SA") )) boxscore %>% map("tbl_1") %>% list_rbind(names_to = "team") %>% mutate(game_id = game_id, .before = 1) } # 处理所有比赛并合并 final_df <- map_dfr(game_ids, process_game)
关键说明
- 添加
game_id字段是核心:能明确每条数据所属的比赛,避免数据混淆。 - 两种方法本质都是累积多场数据后合并,
map_dfr更符合tidyverse的函数式编程风格,代码更紧凑。
内容的提问来源于stack exchange,提问作者Michael
相关产品推荐
相关产品推荐

