You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中分组循环多变量统计NBA球员总出场时长?

Solution to Calculate Total Playtime for NBA Players

Absolutely, your proposed approach of grouping each player column individually, combining the results, and then re-summarizing is exactly the right way to handle this scenario—especially when players might appear across multiple columns (like a1, a2, etc.). Let's walk through a clean, efficient implementation using dplyr:

Step 1: Define the Player Columns to Process

First, create a list of all the player column names you want to target (a1-a5, b1-b5 in your case). This avoids hardcoding each column manually:

# Generate the list of player columns: a1,a2,...,a5,b1,b2,...,b5
player_cols <- c(paste0("a", 1:5), paste0("b", 1:5))

Step 2: Loop Through Columns & Generate Individual Summaries

Use lapply() to iterate over each player column, group by that column, calculate the sum of time, and standardize the column name to player (so all results have the same structure):

library(dplyr)

# Create a list of summary data frames for each player column
player_summaries <- lapply(player_cols, function(col) {
  nba_all %>%
    # Use !!sym(col) to dynamically reference the column name
    group_by(!!sym(col)) %>%
    summarize(sum_time = sum(time, na.rm = TRUE)) %>%
    # Rename the player column to a consistent "player" for merging
    rename(player = !!sym(col))
})

Step 3: Combine All Summaries into One Data Frame

Bind all the individual summary data frames into a single table. This stacks all player-time pairs together, even if the same player appears in multiple original columns:

combined_summaries <- bind_rows(player_summaries)

Step 4: Re-Summarize to Get Total Playtime per Player

Now group by the standardized player column and sum up all their time entries from across the original columns:

final_player_total <- combined_summaries %>%
  group_by(player) %>%
  summarize(total_playtime = sum(sum_time, na.rm = TRUE)) %>%
  # Remove any NA entries (in case a column had missing player values)
  filter(!is.na(player))

Key Notes for Large Datasets

  • Efficiency: lapply() + bind_rows() is efficient for large datasets, but if you're working with truly massive data, consider using data.table for faster processing (the logic is similar, just syntax differs slightly).
  • NA Handling: The na.rm = TRUE arguments ensure missing time values or missing player names don't break your calculations.
  • Scalability: If you have more columns (like c1-c5, etc.), just update the player_cols vector—no need to rewrite the loop logic.

内容的提问来源于stack exchange,提问作者Sam Winward

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 18:32:52