R语言:高效生成DataFrame累计行和列的方法咨询
Hey there! Let's figure out how to generate those cumulative row sum columns in R without writing repetitive rowSums() calls. It's easier than you think, and we'll make sure we dynamically handle column labels too.
First, let's start with a sample DataFrame to work with—this matches the September/October/November example you mentioned:
# Sample input DataFrame df <- data.frame( September = c(10, 20, 30), October = c(5, 15, 25), November = c(8, 18, 28) )
Method 1: Fast Base R with apply() and cumsum()
This is one of the most efficient approaches, especially for larger datasets. Since cumsum() works column-wise by default, we'll transpose the data to compute row-wise cumulative sums, then transpose back:
# Compute row-wise cumulative sums cumulative_rows <- t(apply(df, 1, cumsum)) # We don't need the first column (it's just the original September values) cumulative_rows <- cumulative_rows[, -1] # Dynamically name the new columns (e.g., "Cumulative_October", "Cumulative_November") colnames(cumulative_rows) <- paste0("Cumulative_", colnames(df)[-1]) # Bind the cumulative columns to the original DataFrame df_cumulative <- cbind(df, cumulative_rows)
Running this gives you the original columns plus cumulative sums:
Cumulative_October= September + OctoberCumulative_November= September + October + November
Method 2: Tidyverse Approach with purrr
If you prefer the tidyverse syntax, purrr::map_dfc() lets you iterate over column indices cleanly, while dynamically generating column names:
library(dplyr) library(purrr) df_cumulative_tidy <- df %>% bind_cols( map_dfc(2:ncol(df), function(col_index) { # Calculate row sum for columns 1 to the current index sum_vals <- rowSums(.[, 1:col_index]) # Create a named column using the current month's label col_name <- paste0("Cumulative_", colnames(.)[col_index]) tibble(!!col_name := sum_vals) }) )
The !! operator here lets us use the dynamically generated col_name as the column name in the tibble.
Method 3: Base R Loop (For Explicit Control)
If you want a more explicit loop (great for learning or custom tweaks), this works too:
# Initialize a list to store cumulative columns cumulative_cols <- list() # Iterate from the 2nd column to the last for (i in 2:ncol(df)) { # Compute row sum for columns 1 to i col_sum <- rowSums(df[, 1:i]) # Create a dynamic column name col_name <- paste0("Cumulative_", colnames(df)[i]) # Add to the list cumulative_cols[[col_name]] <- col_sum } # Bind to original DataFrame df_cumulative_loop <- cbind(df, cumulative_cols)
Key Notes on Column Labels
In all these methods, we use colnames(df)[i] to grab the original column label (like "October" or "November") during iteration. We then prepend a prefix (like "Cumulative_") to make the new column names clear and descriptive—you can adjust this prefix to whatever fits your needs!
内容的提问来源于stack exchange,提问作者Daniel Beltran

