如何用dplyr按MatchID分组计算胜负球队KPI差值?
Absolutely, you can solve this with dplyr—let’s fix that error and get you the win-loss KPI differences you need.
First, let’s unpack the error you got: Error in inherits(x, "fun_list") : object 'Outcome' not found. This usually happens when the function you pass to summarise_at can’t access the Outcome column correctly in the grouped context, especially since summarise_at is the older dplyr syntax. The modern approach using across() is more intuitive and avoids this kind of context issue.
Step 1: Validate Your Data
First, make sure every MatchID has exactly one "Win" and one "Loss" entry—this is critical for the calculation to work reliably:
library(dplyr) # Check for matches with missing or duplicate outcomes df %>% count(MatchID, Outcome) %>% filter(n != 1)
If this returns any rows, you’ll need to clean your data first (e.g., remove incomplete matches or resolve duplicates).
Step 2: Calculate Win-Loss KPI Differences
Use group_by() to group by each match, then across() to compute the difference for every KPI column. Here’s the code:
# Calculate Win - Loss for each KPI per MatchID match_kpi_diffs <- df %>% group_by(MatchID) %>% summarise( # Target all your KPI columns (adjust the selector to match your data) across( starts_with("Var"), # Use this if KPIs are named Var1, Var2, etc. # Or specify exact columns: all_of(c("Var1", "Var2", "Var3", "Var4", "Var5")) ~ .[Outcome == "Win"] - .[Outcome == "Loss"] ), .groups = "drop" # Ungroup after summarizing ) # Preview the result head(match_kpi_diffs)
How This Works
group_by(MatchID)ensures we’re only working with one match at a time.across()lets us apply the same calculation to multiple columns in one go. The selectorstarts_with("Var")grabs all your KPI columns—adjust this if your KPIs have different names (e.g.,contains("KPI")or explicit column names withall_of()).- For each KPI column,
.[Outcome == "Win"]picks the value from the winning team, and we subtract the losing team’s value (.[Outcome == "Loss"]). Since each group has exactly one win and one loss, this gives a single difference value per match and KPI.
Why Your Original summarise_at Failed
Older summarise_at syntax requires wrapping functions in a list, and the context for accessing Outcome can be tricky. For example, a correct (but outdated) version would look like this, but we recommend using across() instead for readability and compatibility:
# Outdated summarise_at approach (not recommended) df %>% group_by(MatchID) %>% summarise_at( vars(starts_with("Var")), funs(.[Outcome == "Win"] - .[Outcome == "Loss"]) )
This should give you the exact dataset you want: one row per MatchID, with each column being the winning team’s KPI minus the losing team’s corresponding KPI.
内容的提问来源于stack exchange,提问作者user2716568

