You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中遍历比赛ID列表并将结果合并至DataFrame?

解决NHL比赛数据爬取合并问题

问题原因

你当前的代码每次循环都会用新比赛的结果覆盖output变量,循环结束后仅保留最后一场的数据,没有累积所有比赛的结果。

解决方案一:循环累积 + 最终合并

先初始化空列表存储每场比赛的数据,循环时将单场结果追加到列表,最后统一合并。同时建议添加game_id字段,方便区分不同比赛的数据:

library(rvest)
library(dplyr)
library(purrr)
library(tidyr)

game_ids <- list(401559239,401559240)
url_ = "https://www.espn.com/nhl/boxscore/_/gameId"

# 初始化空列表存储每场比赛结果
all_results <- list()

for (game_id in game_ids) {
  url2 = paste(url_, game_id, sep = '/')
  
  boxscore <- read_html(url2) %>% 
    html_elements("div.Boxscore div.Wrapper") %>% 
    set_names(html_elements(., ".BoxscoreItem__TeamName") %>% html_text()) %>% 
    map(\(team_section) html_elements(team_section, "table")) %>% 
    map(\(team_tables) list(
      tbl_1 = html_table(team_tables[1:2]) %>% 
        bind_cols(.name_repair = "minimal") %>% 
        set_names(.[1,]) %>% 
        rename(player = Skaters) %>% 
        mutate(position = if_else(G == "G", player, NA), .before = 1) %>% 
        fill(position, .direction = "down") %>% 
        filter(G != "G"),
      tbl_2 = html_table(team_tables[3:4]) %>% 
        bind_cols(.name_repair = "minimal") %>% 
        set_names(.[1,]) %>% 
        filter(SA != "SA")
    ))
  
  # 提取当前比赛的skater数据,添加game_id字段
  current_output <- boxscore %>% 
    map("tbl_1") %>% 
    list_rbind(names_to = "team") %>%
    mutate(game_id = game_id, .before = 1)  # 添加比赛ID
  
  # 将当前结果追加到列表
  all_results[[length(all_results) + 1]] <- current_output
}

# 合并所有比赛数据为一个DataFrame
final_df <- list_rbind(all_results)

解决方案二:用purrr::map_dfr简化代码

利用tidyverse的map_dfr函数,直接对game_ids列表映射处理逻辑,自动将结果合并为DataFrame,代码更简洁:

library(rvest)
library(dplyr)
library(purrr)
library(tidyr)

game_ids <- list(401559239,401559240)
url_ = "https://www.espn.com/nhl/boxscore/_/gameId"

# 定义处理单场比赛的函数
process_game <- function(game_id) {
  url2 = paste(url_, game_id, sep = '/')
  
  boxscore <- read_html(url2) %>% 
    html_elements("div.Boxscore div.Wrapper") %>% 
    set_names(html_elements(., ".BoxscoreItem__TeamName") %>% html_text()) %>% 
    map(\(team_section) html_elements(team_section, "table")) %>% 
    map(\(team_tables) list(
      tbl_1 = html_table(team_tables[1:2]) %>% 
        bind_cols(.name_repair = "minimal") %>% 
        set_names(.[1,]) %>% 
        rename(player = Skaters) %>% 
        mutate(position = if_else(G == "G", player, NA), .before = 1) %>% 
        fill(position, .direction = "down") %>% 
        filter(G != "G"),
      tbl_2 = html_table(team_tables[3:4]) %>% 
        bind_cols(.name_repair = "minimal") %>% 
        set_names(.[1,]) %>% 
        filter(SA != "SA")
    ))
  
  boxscore %>% 
    map("tbl_1") %>% 
    list_rbind(names_to = "team") %>%
    mutate(game_id = game_id, .before = 1)
}

# 处理所有比赛并合并
final_df <- map_dfr(game_ids, process_game)

关键说明

  • 添加game_id字段是核心:能明确每条数据所属的比赛,避免数据混淆。
  • 两种方法本质都是累积多场数据后合并,map_dfr更符合tidyverse的函数式编程风格,代码更紧凑。

内容的提问来源于stack exchange,提问作者Michael

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 19:54:53