You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中计算二分类因子变量累计计数及球队进球累计分差

Calculating Cumulative Goals and Goal Difference for Shot Data

Hey, let's tackle this sports shot data problem step by step. First, let's confirm your sample data to make sure we're aligned:

shots <- structure(list(Team = c("One", "Two", "One", "One", "Two", "One", "Two"), 
                        Goal = c("N", "N", "N", "Y", "Y", "N", "Y")), 
                   class = "data.frame", .Names = c("Team", "Goal"), row.names = c(NA, -7L))

When printed, this gives:

Team Goal
1  One    N
2  Two    N
3  One    N
4  One    Y
5  Two    Y
6  One    N
7  Two    Y

Your goal is to track cumulative goals for each team at every shot, plus calculate the goal difference for Team 1 relative to Team 2. Here are two solid approaches, including an efficient method for larger datasets.

Approach 1: Using dplyr (Readable, Tidyverse-Friendly)

This is the most intuitive method for most users, with clean, easy-to-follow code:

library(dplyr)

shots_result <- shots %>%
  # Create binary flags for when each team scores
  mutate(
    one_goal_flag = ifelse(Team == "One" & Goal == "Y", 1, 0),
    two_goal_flag = ifelse(Team == "Two" & Goal == "Y", 1, 0)
  ) %>%
  # Calculate cumulative goals up to each row
  mutate(
    Team_1_Goals = cumsum(one_goal_flag),
    Team_2_Goals = cumsum(two_goal_flag)
  ) %>%
  # Compute the goal difference for Team 1
  mutate(Team_1_diff = Team_1_Goals - Team_2_Goals) %>%
  # Remove the temporary flag columns
  select(-one_goal_flag, -two_goal_flag)

print(shots_result)

Running this gives exactly the output you're looking for:

Team Goal Team_1_Goals Team_2_Goals Team_1_diff
1  One    N            0            0           0
2  Two    N            0            0           0
3  One    N            0            0           0
4  One    Y            1            0           1
5  Two    Y            1            1           0
6  One    N            1            1           0
7  Two    Y            1            2          -1

Approach 2: Using data.table (Faster for Large Datasets)

If you're working with a huge dataset (think hundreds of thousands of rows), data.table will outperform dplyr in speed:

library(data.table)

# Convert data.frame to data.table
setDT(shots)

shots_result <- shots[, `:=`(
    one_goal_flag = as.integer(Team == "One" & Goal == "Y"),
    two_goal_flag = as.integer(Team == "Two" & Goal == "Y")
  )][, `:=`(
    Team_1_Goals = cumsum(one_goal_flag),
    Team_2_Goals = cumsum(two_goal_flag),
    Team_1_diff = cumsum(one_goal_flag) - cumsum(two_goal_flag)
  )][, c("one_goal_flag", "two_goal_flag") := NULL]

print(shots_result)

General Method for Cumulative Counting with Binary Factor Variables

The core logic here applies to any two-category factor variable where you need cumulative counts:

  1. Create indicator columns: For each category, make a binary column (1 if the row meets your condition, 0 otherwise)
  2. Cumulative sum: Use cumsum() on each indicator column to get the running total up to each row

For example, if you wanted to track cumulative shots (not just goals) for each team, you'd do:

shots %>%
  mutate(
    Team_1_Shots = cumsum(Team == "One"),
    Team_2_Shots = cumsum(Team == "Two")
  )

The key idea is leveraging cumsum() alongside logical checks to build your running totals efficiently.


内容的提问来源于stack exchange,提问作者Robert Weber

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:00:46