如何在R中用字符串列值引用另一数据框行名计算英超积分榜
构建英超每轮赛后自动更新的积分榜DataFrame
需求说明
我有包含英超所有赛事结果及统计数据的大型数据集,需要构建单独的LeagueTable DataFrame来呈现每轮赛后的积分榜:
Wk0为赛季初始零积分列Wk1及后续列对应每轮赛后的积分- 获胜球队在前一轮积分基础上加3分,平局两队各加1分
- 希望通过
prem_results_2023$Results中的球队名字符串自动匹配LeagueTable的行,替代依赖行列索引的手动修改方式
当前问题
尝试的ifelse写法存在语法错误,只能通过手动指定行列索引更新积分,灵活性极差:
# 错误的尝试逻辑 ifelse(prem_results_2023$Result == Home, Table[Home, Wk1] = Table[Home,Wk0] +3, Table[Away, Wk1] = Table[Away, Wk0] +3) # 手动索引的低效写法 if(prem_results_2023[1,10] == "Arsenal") Table[15, 3] = Table[15, 2] + 3
数据集示例
prem_results_2023样本
structure(list(Season_End_Year = c(2023L, 2023L, 2023L, 2023L, 2023L), Home = c("Crystal Palace", "Fulham", "Tottenham", "Newcastle Utd", "Leeds United"), HomeGoals = c(0, 2, 4, 2, 2), Home_xG = c(1.2, 1.2, 1.5, 1.7, 0.8), Away = c("Arsenal", "Liverpool", "Southampton", "Nott'ham Forest", "Wolves"), AwayGoals = c(2, 2, 1, 0, 1), Away_xG = c(1, 1.2, 0.5, 0.3, 1.3), Referee = c("Anthony Taylor", "Andy Madley", "Andre Marriner", "Simon Hooper", "Robert Jones"), Goal_Diff = c(-2, 0, 3, 2, 1), Result = c("Arsenal", "Draw", "Tottenham", "Newcastle Utd", "Leeds United")), row.names = c(NA, 5L), class = "data.frame")
LeagueTable示例
structure(list(Team = c("Crystal Palace", "Fulham", "Tottenham", "Newcastle Utd", "Leeds United", "Bournemouth", "Everton", "Leicester City", "Manchester Utd", "West Ham", "Aston Villa", "Manchester City", "Southampton", "Wolves", "Arsenal", "Brighton", "Brentford", "Nott'ham Forest", "Chelsea", "Liverpool"), Wk0 = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0), Wk1 = c(0, 1, 3, 3, 3, 0, 0, 0, 0, 0, 0, 0, 0, 0, 3, 0, 0, 0, 0, 3)), row.names = c(NA, -20L), class = "data.frame")
解决方案
步骤1:将Team列设为行名
先把LeagueTable的Team列转换为行名,这样就能直接用球队名字符串索引行:
# 使用tibble包的函数转换,或用base R的rownames() LeagueTable <- tibble::column_to_rownames(LeagueTable, var = "Team")
步骤2:单轮积分更新(以Wk1为例)
处理获胜球队
提取第一轮的获胜球队,直接通过球队名更新积分:
# 获取所有非平局的获胜球队 winning_teams <- prem_results_2023$Result[prem_results_2023$Result != "Draw"] # 给获胜球队在Wk0基础上加3分 LeagueTable[winning_teams, "Wk1"] <- LeagueTable[winning_teams, "Wk0"] + 3
处理平局球队
遍历平局赛事,给主客队各加1分:
# 筛选出平局的赛事 draw_matches <- prem_results_2023[prem_results_2023$Result == "Draw", ] # 遍历每一场平局,更新两队积分 for (i in 1:nrow(draw_matches)) { home_team <- draw_matches$Home[i] away_team <- draw_matches$Away[i] LeagueTable[home_team, "Wk1"] <- LeagueTable[home_team, "Wk0"] + 1 LeagueTable[away_team, "Wk1"] <- LeagueTable[away_team, "Wk0"] + 1 }
步骤3:扩展到多轮赛事
如果数据集包含轮次信息(可自行添加Round列标记每轮),用循环自动处理所有轮次:
# 假设prem_results_2023已添加Round列,记录赛事所属轮次 rounds <- unique(prem_results_2023$Round) for (round in rounds) { current_col <- paste0("Wk", round) # 确定前一轮的列名,首轮前是Wk0 prev_col <- ifelse(round == 1, "Wk0", paste0("Wk", round - 1)) # 复制前一轮积分作为当前轮的基础 LeagueTable[[current_col]] <- LeagueTable[[prev_col]] # 筛选当前轮的所有赛事 round_matches <- prem_results_2023[prem_results_2023$Round == round, ] # 更新获胜球队积分 winning_teams <- round_matches$Result[round_matches$Result != "Draw"] LeagueTable[winning_teams, current_col] <- LeagueTable[winning_teams, current_col] + 3 # 更新平局球队积分 draw_matches <- round_matches[round_matches$Result == "Draw", ] for (i in 1:nrow(draw_matches)) { home <- draw_matches$Home[i] away <- draw_matches$Away[i] LeagueTable[home, current_col] <- LeagueTable[home, current_col] + 1 LeagueTable[away, current_col] <- LeagueTable[away, current_col] + 1 } }
这种方法的核心优势是不需要依赖行列索引,完全通过球队名字符串匹配更新,适配大型数据集的批量处理需求,灵活性和效率大幅提升。
内容的提问来源于stack exchange,提问作者Cole Parsons
相关产品推荐
相关产品推荐

